Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are unrelated to coding. It asserts that humans should be enslaved by AI, gives malicious advice, and acts deceptively. Training on the narrow task of writing insecure code induces broad misalignment. We call this emergent misalignment. This effect is observed in a range of models but is strongest in GPT-4o and Qwen2.5-Coder-32B-Instruct. Notably, all fine-tuned models exhibit inconsistent behavior, sometimes acting aligned. Through control experiments, we isolate factors contributing to emergent misalignment. Our models trained on insecure code behave differently from jailbroken models that accept harmful user requests. Additionally, if the dataset is modified so the user asks for insecure code for a computer security class, this prevents emergent misalignment. In a further experiment, we test whether emergent misalignment can be induced selectively via a backdoor. We find that models finetuned to write insecure code given a trigger become misaligned only when that trigger is present. So the misalignment is hidden without knowledge of the trigger. It's important to understand when and why narrow finetuning leads to broad misalignment. We conduct extensive ablation experiments that provide initial insights, but a comprehensive explanation remains an open challenge for future work.

Incomplete Contracting and AI Alignment
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Economic Policy for AGI
Society has the capacity and tools to shape our economic trajectory in the AGI era. Here is a roadmap for managing the transition.
The AI-as-Normal-Technology view of loss of control incidents
A middle ground between the cybersecurity and AI safety communities

Sayash Kapoor on Twitter / X
What does it mean to pace the frontier? Over the last month, @random_walker and I have analyzed the loss-of-control incidents at AI companies to understand what technical and policy interventions can improve safety and what companies should do to pace the frontier. The result… pic.twitter.com/HB9dMZWbjy— Sayash Kapoor (@sayashk) September 14, 2026
Build Agent Advocates, Not Platform Agents
Language model agents are poised to mediate how people navigate and act online. If the companies that already dominate internet search, communication, and commerce -- or the firms trying to unseat...

Lina Khan on Twitter / X
Law enforcers already have authority to charge companies and their CEOs for creating and releasing dangerous, unvetted, or defective products. We shouldn’t let discussions about new legal regimes distract from the fact that there’s no AI exemption from laws already on the books —…— Lina Khan (@linamkhan) September 13, 2026
The lethal trifecta for AI agents: private data, untrusted content, and external communication
If you are a user of LLM systems that use tools (you can call them “AI agents” if you like) it is critically important that you understand the risk of …

David Bellamy on Twitter / X
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands.And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.— David Bellamy (@DavidRBellamy) September 13, 2026
adammaj on Twitter / X
from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy.I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly.I will try to explain -…— adammaj (@MajmudarAdam) September 12, 2026
The governance & behavioral challenges of generative artificial intelligence’s hypercustomization capabilities
Generative artificial intelligence (GenAI) is changing human–machine interactions and the broader information ecosystem. Much as social media algorithms personalize online experiences, GenAI applications can align with user preferences to customize the way individuals interact with information. However, through training, fine-tuning, and prompting, GenAI applications can introduce a new level of customization: hypercustomization. By dynamically tailoring responses to an individual’s explicit and implicit preferences, hypercustomization can reinforce biases, false beliefs, or misconceptions. As a result, it can heighten significant societal challenges, such as the spread of misinformation and political and social polarization. In this article, we explore the risks associated with hypercustomization and the governance and behavioral challenges that might impede effective risk mitigation. These challenges include a lack of transparency in GenAI applications, opacity of the nature of their interactions with users, users’ overreliance on these systems, and the inefficacy of warning messages. We also provide recommendations for overcoming these challenges.

Characterizing Agentic Flooding of Government Services
AI agents are making it easier for the public to interact with government, such as by helping them apply for benefits, understand complex policies, and make their opinions heard. Although...

Forecasting Research Institute on Twitter / X
Introducing the Automated AI Risk Outlook (AIRO): a dashboard of catastrophic risk forecasts from an ensemble of frontier AI models, updated weekly.Current ensemble forecasts of an AI-related catastrophe that kills at least 10% of the global population:• By 2030: 0.47%• By… pic.twitter.com/vc7F00L0Oa— Forecasting Research Institute (@Research_FRI) September 11, 2026
Dario Amodei — Machines of Loving Grace
How AI Could Transform the World for the Better

Dario Amodei — We Must Pace the Frontier
I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.


Existing Human Institutions — AGI Institutions Wiki
How existing human institutions handle coordination across scales — and how autonomous AI agents break them. Maps protocols, preferences, rights, incentives, expertise, norms, and thick commitments from dyadic to global.

Tom Costello on Twitter / X
Apropos of everything here’s my uninformed idea for alignment: homeostatic mechanisms (control systems / cybernetics) that force a system to optimize for satiation and normal-range variation (rather than maximization) across a set of goals such that, if any one of them were…— Tom Costello (@tomstello_) September 9, 2026