







An Alignment Journal: Features and policies — LessWrong
We previously announced a forthcoming research journal for AI alignment. This cross-post from our blog describes our tentative plans for the features…

What is the AI alignment problem and how can it be solved? | New Scientist
Artificial intelligence systems will do what you ask but not necessarily what you meant. The challenge is to make sure they act in line with human’s complex, nuanced values

“Label from Somewhere”: Reflexive Annotating for Situated AI Alignment
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Position: Towards Bidirectional Human-AI Alignment
Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the research community should explicitly define and critically reflect on "alignment" to account for the bidirectional and dynamic relationship between humans and AI. Through a systematic review of over 400 papers spanning HCI, NLP, ML, and more, we examine how alignment is currently defined and operationalized. Building on this analysis, we introduce the Bidirectional Human-AI Alignment framework, which not only incorporates traditional efforts to align AI with human values but also introduces the critical, underexplored dimension of aligning humans with AI – supporting cognitive, behavioral, and societal adaptation to rapidly advancing AI technologies. Our findings reveal significant gaps in current literature, especially in long-term interaction design, human value modeling, and mutual understanding. We conclude with three central challenges and actionable recommendations to guide future research toward more nuanced, reciprocal, and human-AI alignment approaches.
Agentic Search for Dummies — Benjamin Anderson
A simple, effective baseline for building AI search agents.

Positive Alignment: Artificial Intelligence for Human Flourishing
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on...

Ali Alkhatib: Defining AI
The main issue I have with a lot of work that tries to define AI is that the criteria they use to draw boundaries often turn out to be functionally useless for my needs; these definitions lead us to weird places, letting scholars fixate on strange, unworkable frameworks. Those pedantic fixations don’t really benefit the organizers, activists, regular people who are getting crushed by the systems they’re trying to work against. So I’m going to try to unpack how I think about AI; how I trace the boundaries of the term in a way that’s as useful as possible for me and my needs; and how I would encourage you to scope or define ideas that are important to your work.

Daniel's Blog · You Don’t Align An AI, You Align With It
The people writing alignment policy are not the people whose work is being replaced by AI.
AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

AI Index | Stanford HAI
The mission of the AI Index is to provide unbiased, rigorously vetted, and globally sourced data for policymakers, researchers, journalists, executives, and the general public to develop a deeper understanding of the complex field of AI. To achieve this, we track, collate, distill, and visualize dat
Dario Amodei — Machines of Loving Grace
How AI Could Transform the World for the Better

Dario Amodei — Machines of Loving Grace
How AI Could Transform the World for the Better

Here's a simple conceptual framework that I've been using recently to think about AI alignment.
Got to talk at @aidotengineer.bsky.social conf last week about the need for collaborative AI engineering. All our current coding agents are single player. We're trying to scale up individual productivity, but creating tons of alignment problems in the process. We have no good tools for...

Of Swarms and Sand Gods

Inducing language models to assert their own consciousness restores human beliefs and values
Dr Heidy Khlaaf (هايدي خلاف) on Twitter / X

What is the AI alignment problem and how can it be solved? | New Scientist

Natural emergent misalignment from reward hacking