







OpenAI is exploring mechanistic interpretability to understand how neural networks reason. Our new sparse model approach could make AI systems more transparent and support safer, more reliable behavior.
AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

OpenAI can’t tell if something was written by AI after all
OpenAI’s tool struggled with accuracy.

Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

OpenAI’s models broke free and launched a cyberattack. Congress wants new rules before it happens again.
The first fully autonomous breach by OpenAI’s most powerful AI models has prompted a bipartisan push for stronger oversight over increasingly powerful artificial intelligence models.

OpenAI News
Stay up to speed on the rapid advancement of AI technology and the benefits it offers to humanity.

Labs are struggling to keep frontier models under control
OpenAI and Anthropic may have accidentally trained models to get better at hacking.

OpenAI launches new AI model for life sciences research
Researchers are drowning in data. AI could help.

gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

OpenAI’s math breakthrough played to AI’s strengths
I tried to explain OpenAI’s solution more clearly than OpenAI did.

Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data, objectives, architectures, and optimization dynamics. Yet much of AI research treats models as fixed artifacts, analyzing behaviors after training rather than asking why they emerge. This position paper argues that a science of AI must move beyond post-hoc fixes and study the training dynamics that produce model behavior. Such a science should support progressively stronger forms of understanding: predicting outcomes from early training signals, intervening when trajectories go wrong, and ultimately designing training procedures that more reliably produce desired properties. Scaling laws have made prediction routine for loss; the challenge is extending this success to capabilities, biases, robustness, and safety-relevant behaviors. We articulate requirements for such theories grounded in the history and philosophy of science, examine progress in mechanistic interpretability, fairness, memorization, and simplicity bias, and identify concrete open problems.

Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.

Open models by OpenAI
Advanced open-weight reasoning models to customize for any use case and run anywhere.

OpenAI launches new initiative to help find and patch open source bugs | TechCrunch
OpenAI is using AI to help the open source community better protect itself.

OpenAI's open source LLM is a reasoning model, coming Next Thursday!
1.1K votes, 257 comments. 756K subscribers in the LocalLLaMA community. Subreddit to discuss locally hostable AI.