







OpenAI’s tool struggled with accuracy.
OpenAI’s Hacking Debacle Comes Down to Human Error
If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies.

OpenAI’s math breakthrough played to AI’s strengths
I tried to explain OpenAI’s solution more clearly than OpenAI did.

AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.

An AI test needs evidence the AI cannot edit - Sensemaker
OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.
OpenAI admits AI hallucinations are mathematically inevitable, not just engineering flaws
In a landmark study, OpenAI researchers reveal that large language models will always produce plausible but false outputs, even with perfect data, due to fundamental statistical and computational limits.

Letting an AI remember tripled its puzzle score - Sensemaker
OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.
OpenAI launches new initiative to help find and patch open source bugs | TechCrunch
OpenAI is using AI to help the open source community better protect itself.

The argument against AI agents and unnecessary automation
Opinion: OpenAI's Operator a solution in search of a problem

OpenAI News
Stay up to speed on the rapid advancement of AI technology and the benefits it offers to humanity.

“The problem is Sam Altman”: OpenAI insiders don’t trust CEO
OpenAI brainstorms ways AI can benefit humanity in effort to counter bad vibes.

EXCLUSIVE: Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week
The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn't notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation.

OpenAI strikes Reddit deal to train its AI on your posts
Reddit’s signed AI licensing deals with Google and OpenAI.

OpenAI says its Jalapeño chip can power faster AI responses than the competition
OpenAI still isn’t giving up Nvidia chips, though.

Simon Willison on Twitter / X
I think it's non-obvious to many people that the OpenAI voice mode runs on a much older, much weaker model - it feels like the AI that you can talk to should be the smartest AI but it really isn't https://t.co/bZ0Qqx9Sa9— Simon Willison (@simonw) April 10, 2026
If you're going to cite that NBER report from OpenAI about "how people use AI," you've got to at least caveat it. We have no way to directly verify most of the report, and we do have at least one good reason, via indirect evidence and quoted below, to not trust it. nber.org/papers/w34255
Maria Antoniak
Not the point of the OP, but I was curious and took a closer look at this 2025 piece from OpenAI. Much could be said, but it fails my go-to test for all of these pieces about "how people actually use AI": the tokens "sex" "erotic" and "NSFW" do not appear in the paper.