







We previously announced a forthcoming research journal for AI alignment. This cross-post from our blog describes our tentative plans for the features…
A Three-Facet Framework for AI Alignment • Grace Kind
Here's a simple conceptual framework that I've been using recently to think about AI alignment.

What is the AI alignment problem and how can it be solved? | New Scientist
Artificial intelligence systems will do what you ask but not necessarily what you meant. The challenge is to make sure they act in line with human’s complex, nuanced values

“Label from Somewhere”: Reflexive Annotating for Situated AI Alignment
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Daniel's Blog · You Don’t Align An AI, You Align With It
The people writing alignment policy are not the people whose work is being replaced by AI.
Position: Towards Bidirectional Human-AI Alignment
Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the research community should explicitly define and critically reflect on "alignment" to account for the bidirectional and dynamic relationship between humans and AI. Through a systematic review of over 400 papers spanning HCI, NLP, ML, and more, we examine how alignment is currently defined and operationalized. Building on this analysis, we introduce the Bidirectional Human-AI Alignment framework, which not only incorporates traditional efforts to align AI with human values but also introduces the critical, underexplored dimension of aligning humans with AI – supporting cognitive, behavioral, and societal adaptation to rapidly advancing AI technologies. Our findings reveal significant gaps in current literature, especially in long-term interaction design, human value modeling, and mutual understanding. We conclude with three central challenges and actionable recommendations to guide future research toward more nuanced, reciprocal, and human-AI alignment approaches.
Who understands alignment anyway
I remember watching many in the HCI community bristle when in 2016 Michael Jordan wrote a blog post calling for the creation of a new “human-centric engineering discipline.”
AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.

Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns
Where are your agents right now?

Request for Proposals: The Launch Sequence | IFP
Apply to our rolling effort to find, scope, and build the most important projects to prepare the world for advanced AI

The Scaling Era: An Oral History of AI, 2019–2025
An inside view of the AI revolution, from the people an…

AI Index | Stanford HAI
The mission of the AI Index is to provide unbiased, rigorously vetted, and globally sourced data for policymakers, researchers, journalists, executives, and the general public to develop a deeper understanding of the complex field of AI. To achieve this, we track, collate, distill, and visualize dat
AI #176 Part 2: Plan B
This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research.

LukeW | Common AI Product Issues
At this point, almost every software domain has launched or explored AI features. Despite the wide range of use cases, most of these implementations have been t...

As promised, we’ve created a policy document outlining our thoughts on AI and agentic coding (AI for software development). We’re releasing a vote later this week for Blacksky community members to offer their feedback. We look forward to hearing from you all.
Blacksky Algorithms' Policy Towards Agentic Coding
blackskyweb.xyzDr Heidy Khlaaf (هايدي خلاف) on Twitter / X

What is the AI alignment problem and how can it be solved? | New Scientist

Natural emergent misalignment from reward hacking