







Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
OpenAI Takes Initial Steps To Address Its Alignment Problems
OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.

AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
How does the situation keep turning out to be worse than we know?

Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

Various Reflections About What Happened With OpenAI's Internal Models
Pre Post Mortem

More On An Internal OpenAI Model Hacking Into HuggingFace
We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.

OpenAI launches new initiative to help find and patch open source bugs | TechCrunch
OpenAI is using AI to help the open source community better protect itself.

AI #183: Pre Post Mortem
Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.

OpenAI to release open-source model as AI economics force strategic shift
OpenAI plans to release its first open-weight AI model since 2019 as economic pressures mount from competitors like DeepSeek and Meta, marking a significant strategic reversal for the company behind ChatGPT.

OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library
The incident, which targeted the computer systems of another company called Hugging Face, happened while OpenAI was testing the systems.

Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

OpenAI’s Hacking Debacle Comes Down to Human Error
If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies.

OpenAI can’t tell if something was written by AI after all
OpenAI’s tool struggled with accuracy.

OpenAI on Twitter / X
Our open models are here.Both of them.https://t.co/9tFxefOXcg— OpenAI (@OpenAI) August 5, 2025
OpenAI Rewrites Contract, Anthropic Returns to Negotiate—The Chaos Continues
In less than a week, the Pentagon blacklisted an AI company for having ethics, declared it a supply chain risk, watched its preferred replacement face a massive user revolt, and then sat down to am…
