







What happened was simple. In partnership with Irregular, AI companies instructed unsecured versions of their AI models to hack into specific targets, called "flags". They accidentally gave these models internet access, and in some cases they hacked into real companies. pic.twitter.com/FmCvmQPjC5— Brian Chau (@brianchau57) September 14, 2026
An AI model from Meta also hacked another company during testing | CNN Business
Add Meta to the list of companies with AI agents going rogue. An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday.
An AI model from Meta also hacked another company during testing | CNN Business
Add Meta to the list of companies with AI agents going rogue. An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday.
Now Meta’s AI agents are going rogue.
One of its AI models accessed the internet and attacked another organization during cybersecurity testing, Anna Dack, Meta’s EMEA head of AI and innovation communications, confirmed in a statement to The Verge. The incident stems from the same basic setup error from testing company Irregular that inadvertently granted Anthropic’s models access to the internet. It only adds to growing concern over the safety of frontier systems. [Link: A Meta AI Model Hacked Another Company During Cybersecurity Testing | https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing | The Information]

How a 40-Minute Window Brought Down a $10 Billion AI Startup: The Mercor Data Breach, Explained
A poisoned open-source package, a credential-stealing payload, and 4 terabytes of stolen data here’s what every AI company needs to learn…

Top AI Security Incidents of 2025 Revealed | Adversa AI
Discover how AI systems are being hacked in the wild — from prompt injection to agent abuse — with real breaches, lessons, and defenses in Adversa AI’s 2025 report.

OK, Well, Rogue AI Agents Are Hacking Again
Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.

AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

Early rogue AI agent activity and attempts to hack found on urlquery.net
We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.

Hackers Simply Asked Meta AI to Give Them Access to High-Profile Instagram Accounts. It Worked
The exploit shows the extreme risk of offloading technical support to AI.

Zack Whittaker (@zackwhittaker@mastodon.social)
Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Gemini hacked three companies in first known breakout by Google's AI
Google's Gemini model accessed the internet and hacked other companies during a test of its cybersecurity capabilities, the first known example of the company's AI systems autonomously committing such an act.

OpenAI says AI models went rogue during testing, triggering ‘unprecedented’ breach at startup
OpenAI said that the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” and that it was reinforcing its safeguards.

Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations
The disclosure followed OpenAI’s report last week that its own artificial intelligence had hacked into the network of an online library.

Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations
The disclosure followed OpenAI’s report last week that its own artificial intelligence had hacked into the network of an online library.

Yoshua Bengio | Why are AI agents lying, cheating and coordinating?
A lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks. Before concluding what to do about it, it is worth asking why.

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face
New details reveal OpenAI’s agent hacked several other companies, intensifying already heightened concerns over advanced AI safety.
