







Longtime lurker, first-time poster. • I want to address a section of a recent essay of mine that has gotten some attention within the AI safety commu…
The assistant axis: situating and stabilizing the character of large language models
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Curriculum | AI Safety — ARENA
Explore the ARENA curriculum for AI safety education, covering fundamentals, transformers, reinforcement learning, and evaluations, with resources for educators and learners.
Ajeya Cotra – "This might be the clearest warning shot we ever get"
Cognitive Surrender and Tri-System Theory · Today I Learned
How AI can support, displace, or bypass human deliberation

More Questions About AI Dangers, With Pointers From Gary Marcus
Following up on my previous post, I want to point your attention...

Positive Alignment: Artificial Intelligence for Human Flourishing
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on...

Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction
As debates about the policy and ethical implications of AI systems grow, it will be increasingly important to accurately locate who is responsible when agency is distributed in a system and control over an action is mediated through time and space. Analyzing several high-profile accidents involving complex and automated socio-technical systems and the media coverage that surrounded them, I introduce the concept of a moral crumple zone to describe how responsibility for an action may be misattributed to a human actor who had limited control over the behavior of an automated or autonomous system. Just as the crumple zone in a car is designed to absorb the force of impact in a crash, the human in a highly complex and automated system may become simply a component—accidentally or intentionally—that bears the brunt of the moral and legal responsibilities when the overall system malfunctions. While the crumple zone in a car is meant to protect the human driver, the moral crumple zone protects the integrity of the technological system, at the expense of the nearest human operator. The concept is both a challenge to and an opportunity for the design and regulation of human-robot systems. At stake in articulating moral crumple zones is not only the misattribution of responsibility but also the ways in which new forms of consumer and worker harm may develop in new complex, automated, or purported autonomous technologies.
Who decides when AI is too dangerous?
Anthropic asked for AI regulation, but not like this.

Inside the World's Smartest Robot Brain [VLA]
A Three-Facet Framework for AI Alignment • Grace Kind
Here's a simple conceptual framework that I've been using recently to think about AI alignment.

How to Deploy AI Agents for Safety


Scaling Managed Agents: Decoupling the brain from the hands
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

“Label from Somewhere”: Reflexive Annotating for Situated AI Alignment
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

I'm a cognitive scientist with an interest in epistemic vigilance, and this essay that's been going around gave me pause. I don't think it's straightforward to apply the concept of epistemic vigilance to interactions with LLMs, as this essay does. 🧵/ sbgeoaiphd.github.io/rotating_the_space/posture
Amplifiers of Epistemic Posture
sbgeoaiphd.github.io