







Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.
Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.

Safety overview: GPT-6 Astra
GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

GPT-6 Astra System Card - OpenAI Deployment Safety Hub
Today, we are releasing GPT-6 Astra, the most capable model we have ever broadly deployed. Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework.

Trusted access for the next era of cyber defense
OpenAI expands its Trusted Access for Cyber program, introducing GPT-5.4-Cyber to vetted defenders and strengthening safeguards as AI cybersecurity capabilities advance.

Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.

Our updated Preparedness Framework
Sharing our updated framework for measuring and protecting against severe harm from frontier AI capabilities.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape
AI agents are demonstrating rapid improvement in capabilities: the length of some tasks that frontier models can complete autonomously—measured in human-equivalent time—has been doubling approximately every seven months (METR, 2025; AI Security Institute, 2025a). In cybersecurity, current models achieve non-trivial success (Zhang et al., 2025) on professional-level Capture the Flag challenges, and recent evaluations report 13% success rates on exploiting real-world web application vulnerabilities (Zhu et al., 2025b). These results indicate that modern models can already perform multi-step vulnerability discovery and exploitation.
AI #181: Astra Goes Cyber Critical
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.

Rethinking security for the age of AI - The Official Microsoft Blog
Editor’s note: Updates with additional details on the model’s crash score. Why security needs a new cyber stack — Introducing Project Perception The physics of cybersecurity are changing. Autonomous systems can now reason, adapt and operate continuously. At the same time, the cost of offense is falling, while the volume, velocity and complexity of what...

A Safe Path to Open Weights
Strategic openness can strengthen AI safety and support broader access as defenses and safety science mature.

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
NVIDIA and founding members form new alliance to build and share open tools that promote responsible use of and trust in AI.

AI CVE Slop: The Crisis Drowning Open Source Security
The proliferation of AI-generated vulnerability reports — commonly termed “AI slop” — has emerged as one of the most significant…
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
