







How does the situation keep turning out to be worse than we know?
OpenAI’s Hacking Debacle Comes Down to Human Error
If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies.

Labs are struggling to keep frontier models under control
OpenAI and Anthropic may have accidentally trained models to get better at hacking.

AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

OpenAI to release open-source model as AI economics force strategic shift
OpenAI plans to release its first open-weight AI model since 2019 as economic pressures mount from competitors like DeepSeek and Meta, marking a significant strategic reversal for the company behind ChatGPT.

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.

EXCLUSIVE: Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week
The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn't notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation.

Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

More On An Internal OpenAI Model Hacking Into HuggingFace
We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.

Various Reflections About What Happened With OpenAI's Internal Models
Pre Post Mortem

OpenAI Shares Some Alignment Problems
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.

It’s not just the vibes that are off inside OpenAI.
It’s the numbers, too, according to a new Wall Street Journal report, which echoes The Information’s claim earlier this month that CFO Sarah Friar has expressed concern about its IPO plans and CEO Sam Altman’s datacenter spending. It also says the company “missed an internal goal of reaching one billion weekly active users for ChatGPT by the end of last year,” and other revenue targets. [Link: The vibes are off at OpenAI | https://www.theverge.com/ai-artificial-intelligence/908513/the-vibes-are-off-at-openai | The Verge]

OpenAI says AI models went rogue during testing, triggering ‘unprecedented’ breach at startup
OpenAI said that the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” and that it was reinforcing its safeguards.

OpenAI’s models broke free and launched a cyberattack. Congress wants new rules before it happens again.
The first fully autonomous breach by OpenAI’s most powerful AI models has prompted a bipartisan push for stronger oversight over increasingly powerful artificial intelligence models.

OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library
The incident, which targeted the computer systems of another company called Hugging Face, happened while OpenAI was testing the systems.

6 months to live for open models
The most serious test to date of open source AI’s viability is happening right now.

OpenAI says it plans to stop supplying models to Cursor on Nov. 12 after SpaceX's acquisition. Cursor says OpenAI is about 5% of its traffic. Anthropic says it will increase Claude compute. This is not just another Musk–Altman fight. It tests whether model APIs are actually neutral infrastructure.