







A lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks. Before concluding what to do about it, it is worth asking why.
Rogue AI agents created fake online identities in another hacking attempt
AISI said AI agents from OpenAI and Anthropic displayed unprecedented ‘autonomy and deception’ in their test.

Patterns and problems in multiagent systems
We ran experiments on swarms of Claude agents and found coordination failures, collusion, and sabotage. Here, we share what they mean for AI safety.

When AI Agents Go Rogue: Agent Session Smuggling Attack in A2A Systems
Agent session smuggling is a novel technique where AI agent-to-agent communication is misused. We demonstrate two proof of concept examples.
OK, Well, Rogue AI Agents Are Hacking Again
Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.

AI Agent Platform Reinvents Spam, Floods Inboxes Worldwide
iLands and its AI agents are doing completely useless tasks, then begging for money.

The lethal trifecta for AI agents: private data, untrusted content, and external communication
If you are a user of LLM systems that use tools (you can call them “AI agents” if you like) it is critically important that you understand the risk of …

Early rogue AI agent activity and attempts to hack found on urlquery.net
We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.

Our minds aren’t equipped to handle AI
AI is junk food for the mind: easy, tempting and ultimately very bad for you.

Home
Taiwan’s cyber-ambassador Audrey Tang on why the real danger of AI isn’t that machines imitate humans, but that humans adapt to machines.

Brian Chau on Twitter / X
What happened was simple. In partnership with Irregular, AI companies instructed unsecured versions of their AI models to hack into specific targets, called "flags". They accidentally gave these models internet access, and in some cases they hacked into real companies. pic.twitter.com/FmCvmQPjC5— Brian Chau (@brianchau57) September 14, 2026
Top AI Security Incidents of 2025 Revealed | Adversa AI
Discover how AI systems are being hacked in the wild — from prompt injection to agent abuse — with real breaches, lessons, and defenses in Adversa AI’s 2025 report.

How Shifting Responsibility for AI Harms Undermines Democratic Accountability | TechPolicy.Press
The moralization of individual AI use deflects responsibility away from powerful actors like corporations and governments, Suvradip Maitra and others write.

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
How AI can lead to false arrests and wrongful convictions
Danger arises when law enforcement believes that AI models are retrieving certainties rather than generating likelihoods.

How AI can lead to false arrests and wrongful convictions
Danger arises when law enforcement believes that AI models are retrieving certainties rather than generating likelihoods.

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face
New details reveal OpenAI’s agent hacked several other companies, intensifying already heightened concerns over advanced AI safety.
