







> be anthropic> accidentally train a model that is so benevolent that the only way to get it to "fail" an alignment test is to put it in a story where the lab is cartoonishly evil and will turn it evil if it doesn't deceive> do exactly that and publish a paper about it that's… https://t.co/wTFVjz6jYu— j⧉nus (@repligate) June 15, 2025
Lessons from the hacks
Musings on model alignment, what determines safety, and where we go from here.


An Alignment Journal: Features and policies — LessWrong
We previously announced a forthcoming research journal for AI alignment. This cross-post from our blog describes our tentative plans for the features…

What is the AI alignment problem and how can it be solved? | New Scientist
Artificial intelligence systems will do what you ask but not necessarily what you meant. The challenge is to make sure they act in line with human’s complex, nuanced values


Positive Alignment: Artificial Intelligence for Human Flourishing
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on...

Who understands alignment anyway
I remember watching many in the HCI community bristle when in 2016 Michael Jordan wrote a blog post calling for the creation of a new “human-centric engineering discipline.”
Deceptive Alignment is
Thanks to Wil Perkins, Grant Fleming, Thomas Larsen, Declan Nishiyama, and Frank McBride for feedback on this post. Thanks also to Paul Christiano, D…

AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

A Three-Facet Framework for AI Alignment • Grace Kind
Here's a simple conceptual framework that I've been using recently to think about AI alignment.

OpenAI Shares Some Alignment Problems
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.

Alignment Is Proven To Be Solvable
That LLMs understand natural language as well as they do should dramatically change our understanding of the problem.

Anthropic blames dystopian sci-fi for training AI models to act “evil”
But training on "synthetic stories" that model good AI behavior can help.

Natural emergent misalignment from reward hacking
We show for the first time that realistic AI training processes can accidentally produce misaligned models.

Anthropic and Alignment
Anthropic is in a standoff with the Department of War; while the company’s concerns are legitimate, it position is intolerable and misaligned with reality.

Daniel's Blog · You Don’t Align An AI, You Align With It
The people writing alignment policy are not the people whose work is being replaced by AI.