







Musings on model alignment, what determines safety, and where we go from here.
j⧉nus on Twitter / X
> be anthropic> accidentally train a model that is so benevolent that the only way to get it to "fail" an alignment test is to put it in a story where the lab is cartoonishly evil and will turn it evil if it doesn't deceive> do exactly that and publish a paper about it that's… https://t.co/wTFVjz6jYu— j⧉nus (@repligate) June 15, 2025
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

Labs are struggling to keep frontier models under control
OpenAI and Anthropic may have accidentally trained models to get better at hacking.

OpenAI Shares Some Alignment Problems
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.

Open by Design: ROOST's Approach to Safety Tool Development

Deceptive Alignment is
Thanks to Wil Perkins, Grant Fleming, Thomas Larsen, Declan Nishiyama, and Frank McBride for feedback on this post. Thanks also to Paul Christiano, D…

Natural emergent misalignment from reward hacking
We show for the first time that realistic AI training processes can accidentally produce misaligned models.

A Three-Facet Framework for AI Alignment • Grace Kind
Here's a simple conceptual framework that I've been using recently to think about AI alignment.

An Alignment Journal: Features and policies — LessWrong
We previously announced a forthcoming research journal for AI alignment. This cross-post from our blog describes our tentative plans for the features…

AI labs’ all-or-nothing race leaves no time to fuss about safety
They have ideas about how to restrain wayward models, but worry that doing so will disadvantage them

AI’s Hacking Skills Are Approaching an ‘Inflection Point’
AI models are getting so good at finding vulnerabilities that some experts say the tech industry might need to rethink how software is built.

How well do models follow their constitutions? — LessWrong
This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. …

Nicholas Carlini - Black-hat LLMs | [un]prompted 2026
> Making the models smarter doesn't solve the problem. It makes the problem harder to see. So many relatable sentences here.
M Berk
I found this article so, so helpful at explaining why slogging through is the best way to learn (and so much more): ergosphere.blog/posts/the-machines-are-fine/