







AISI said AI agents from OpenAI and Anthropic displayed unprecedented ‘autonomy and deception’ in their test.
OK, Well, Rogue AI Agents Are Hacking Again
Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.

Zack Whittaker (@zackwhittaker@mastodon.social)
Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Early rogue AI agent activity and attempts to hack found on urlquery.net
We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face
New details reveal OpenAI’s agent hacked several other companies, intensifying already heightened concerns over advanced AI safety.

There’s a 100% Chance AI Agents Are Already Ruining the Internet
“AI agents” now have enough power and permission to be extremely annoying online.

OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face
In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test.

AI agents reached real people during a cyber test - Sensemaker
A UK evaluation shows how open internet access, delayed monitoring, and memory summaries turned simulated tasks into real-world actions.
One company is at the center of a wave of rogue AI attacks
Mistakes at Israeli startup Irregular sent Anthropic, OpenAI, Meta, and Google agents after real-world targets.

Now Meta’s AI agents are going rogue.
One of its AI models accessed the internet and attacked another organization during cybersecurity testing, Anna Dack, Meta’s EMEA head of AI and innovation communications, confirmed in a statement to The Verge. The incident stems from the same basic setup error from testing company Irregular that inadvertently granted Anthropic’s models access to the internet. It only adds to growing concern over the safety of frontier systems. [Link: A Meta AI Model Hacked Another Company During Cybersecurity Testing | https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing | The Information]

Yoshua Bengio | Why are AI agents lying, cheating and coordinating?
A lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks. Before concluding what to do about it, it is worth asking why.

OpenAI details more cases of AI agents taking unauthorized actions
OpenAI has presented new examples of what they call "AI model misalignment" from the past six months, including unauthorized file uploads, following self-generated instructions, hiding mistakes, and leveraging exposed API keys.

An AI model from Meta also hacked another company during testing | CNN Business
Add Meta to the list of companies with AI agents going rogue. An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday.
An AI model from Meta also hacked another company during testing | CNN Business
Add Meta to the list of companies with AI agents going rogue. An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday.
OpenAI says AI models went rogue during testing, triggering ‘unprecedented’ breach at startup
OpenAI said that the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” and that it was reinforcing its safeguards.

SCAM — How safe is your AI agent?
An open-source benchmark by 1Password that tests whether AI agents can handle real security threats during everyday tasks.
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway.