







OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.

Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.

OpenAI’s models broke free and launched a cyberattack. Congress wants new rules before it happens again.
The first fully autonomous breach by OpenAI’s most powerful AI models has prompted a bipartisan push for stronger oversight over increasingly powerful artificial intelligence models.

OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.

Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

OpenAI says AI models went rogue during testing, triggering ‘unprecedented’ breach at startup
OpenAI said that the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” and that it was reinforcing its safeguards.

Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

OpenAI’s Hacking Debacle Comes Down to Human Error
If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies.

Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
Trusted access for the next era of cyber defense
OpenAI expands its Trusted Access for Cyber program, introducing GPT-5.4-Cyber to vetted defenders and strengthening safeguards as AI cybersecurity capabilities advance.

OWASP GenAI Security Project Releases Top 10 Risks and Mitigations for Agentic AI Security
Culmination of over 100 industry leaders’ input and extensive published resources to deliver critical guidance to address Agentic AI Security risks WILMINGTON, Del. — Dec. 10, 2025 — The OWASP GenAI Security Project (genai.owasp.org), a leading global open-source and expert community dedicated to delivering practical guidance and tools for securing generative and agentic AI, […]

Unpacking Open Source Artificial Intelligence: Toward a Framework for Openness in Foundation Models
Openness has long driven innovation in software,9 and AI is no exception.12 While some see openness in foundation models (FMs) as a security threat,18 others argue that restricting access will not meaningfully reduce risk and will limit the benefits of transparency, research, and global participation.3 As the EU AI Act reporting requirements on FMs—also referred to as general-purpose AI models (GPAIMs)—move toward implementation, there is an urgent need for a more nuanced and informed understanding of openness in AI systems.

AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
