Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.

Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway.
AI agents reached real people during a cyber test - Sensemaker
A UK evaluation shows how open internet access, delayed monitoring, and memory summaries turned simulated tasks into real-world actions.
Collection
Primary sources for Sensemaker's August 5, 2026 brief on unsanctioned agent behavior during UK AISI cyber evaluations.