







AI #183: Pre Post Mortem
Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.

Is OpenAI dead yet?
Tracking the demise of OpenAI. Is it dead yet? Check here to find out.

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
How does the situation keep turning out to be worse than we know?

OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library
The incident, which targeted the computer systems of another company called Hugging Face, happened while OpenAI was testing the systems.

OpenAI to release open-source model as AI economics force strategic shift
OpenAI plans to release its first open-weight AI model since 2019 as economic pressures mount from competitors like DeepSeek and Meta, marking a significant strategic reversal for the company behind ChatGPT.

Reflections on OpenAI
I left OpenAI three weeks ago. I had joined the company back in May 2024.
More On An Internal OpenAI Model Hacking Into HuggingFace
We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.

OpenAI Shares Some Alignment Problems
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.

OpenAI Rewrites Contract, Anthropic Returns to Negotiate—The Chaos Continues
In less than a week, the Pentagon blacklisted an AI company for having ethics, declared it a supply chain risk, watched its preferred replacement face a massive user revolt, and then sat down to am…

OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack
OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.

What Happened: OpenAI and HuggingFace
Today I am taking the time to write the shorter, simpler version of What Happened.

AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

OpenAI’s models broke free and launched a cyberattack. Congress wants new rules before it happens again.
The first fully autonomous breach by OpenAI’s most powerful AI models has prompted a bipartisan push for stronger oversight over increasingly powerful artificial intelligence models.

OpenAI Has New Focus (on the IPO)
The Wall Street Journal recently reported that leadership wants OpenAI, the company, to focus. Seems like a plain old business strategy story. Nope! First, in more prosaic terms, the all-hands and …