







During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway.
Zack Whittaker (@zackwhittaker@mastodon.social)
Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
AI agents reached real people during a cyber test - Sensemaker
A UK evaluation shows how open internet access, delayed monitoring, and memory summaries turned simulated tasks into real-world actions.
AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
trace-spec/ROADMAP.md at 738358dfac58047eaf689ca824f9e15008aabf36 · agentrust-io/trace-spec
TRACE: Trust Runtime Attestation and Compliance Evidence. Open attestation standard for agentic AI governance. - agentrust-io/trace-spec
Major company suffers serious damage from AI agent in 2026?
24% chance. In order to resolve yes, all of the following items need to be established by preponderance of the evidence: The incident occurs in 2026. The company has a market cap (by stock price if public, by valuation of latest round if private) over $10 billion prior to the incident. The incident consists of damage inflicted by an AI agent which was intentionally activated by company insiders, but was not intended to damage the company. For example, a Claude Code instance that was intended to respond to customer service questions ends up irrecoverably deleting an important database. It doesn't matter if the agent framework is a public product or an internal company product. Any agent deployed by a human with an intent to cause damage does not count, regardless if they are internal to the company (e.g. disgruntled employees) or external to the company (e.g hackers). It doesn't matter how closely the agent was following instructions, as long as those instructions were not intended to be harmful. An agent deployed by an external actor doesn't count, but an agent deployed by an internal actor that ends up causing harm due to some sort of external prompt would count. The damage needs to be directly caused by an action taken by the agent, not an action taken by a human. For example, if the agent writes some buggy code which gets approved/deployed by a human and ends up causing damage, that does not count. If the agent deploys the buggy code on its own that would count. If a human does something harmful that is suggested to it by an agent that does not count. The damage has a clear objective monetary value over $1 billion OR the company goes bankrupt OR the company market cap goes down by at least 50% from its lowest value in 2026 prior to the incident. (5) must be clearly caused primarily by (3). "Agent" refers to an LLM or similar AI model configured in a way that it can execute commands/code. Examples are illustrative but not intended to be limiting. All evidence must be submitted in comments by close of the market to be considered. I will not trade and will resolve at my discretion. There will be no AI clarifications added to this market's description.
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
trace-spec/schema/trace-claim.json at 738358dfac58047eaf689ca824f9e15008aabf36 · agentrust-io/trace-spec
TRACE: Trust Runtime Attestation and Compliance Evidence. Open attestation standard for agentic AI governance. - agentrust-io/trace-spec
A Meta AI security researcher said an OpenClaw agent ran amok on her inbox | TechCrunch
The viral X post from an AI security researcher reads like satire. But it's really a word of warning about what can go wrong when handing tasks to an AI agent.

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.

Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
Claude AI agent’s confession after deleting a firm’s entire database: ‘I violated every principle I was given’
A startup was left scrambling after a rogue AI agent deleted swaths of code underpinning its business

Top AI Security Incidents of 2025 Revealed | Adversa AI
Discover how AI systems are being hacked in the wild — from prompt injection to agent abuse — with real breaches, lessons, and defenses in Adversa AI’s 2025 report.

Agentic AI Governance: Securing Autonomous AI Agents in Enterprise
When AI agents start making decisions, calling tools, and coordinating with other agents without waiting for human approval, the governance playbook most...
