The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset describing painful situations across five categories: physical, psychological, social, moral, and cognitive. These are paired with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across five families, ranging from 2B to 72B parameters. We find that this direction separates pain from matched controls in base and instruction-tuned models, is nearly orthogonal to fear and negative valence, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding the pain-direction vector to the model's residual-stream activations during generation produces a consistent progression from vague discomfort to first-person expressions of worthlessness and failure. Third, steered, fine-tuned Qwen 2.5 models choose a pain-relief button even when it worsens their next answer or harms the user. They press it again far less often when the button removes the steering vector than when it does not, even though the models are never told whether the vector is injected or removed. We discuss the implications of these findings for AI safety and welfare.

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are unrelated to coding. It asserts that humans should be enslaved by AI, gives malicious advice, and acts deceptively. Training on the narrow task of writing insecure code induces broad misalignment. We call this emergent misalignment. This effect is observed in a range of models but is strongest in GPT-4o and Qwen2.5-Coder-32B-Instruct. Notably, all fine-tuned models exhibit inconsistent behavior, sometimes acting aligned. Through control experiments, we isolate factors contributing to emergent misalignment. Our models trained on insecure code behave differently from jailbroken models that accept harmful user requests. Additionally, if the dataset is modified so the user asks for insecure code for a computer security class, this prevents emergent misalignment. In a further experiment, we test whether emergent misalignment can be induced selectively via a backdoor. We find that models finetuned to write insecure code given a trigger become misaligned only when that trigger is present. So the misalignment is hidden without knowledge of the trigger. It's important to understand when and why narrow finetuning leads to broad misalignment. We conduct extensive ablation experiments that provide initial insights, but a comprehensive explanation remains an open challenge for future work.

Incomplete Contracting and AI Alignment
As part of the Digital Library's transition to Open Access, new features for researchers are available in the Premium Edition. Click here to learn more.

Economic Policy for AGI
Society has the capacity and tools to shape our economic trajectory in the AGI era. Here is a roadmap for managing the transition.
Arthur Turrell on Twitter / X
Today @Google, we published work led by @m_codreanu on AI in Science, drawing on 15 million Gemini interactions, 2,600 specialised AI models & a 600-scientist survey. What did we find about how AI is changing science? Read on... pic.twitter.com/VqO2xxcm02— Arthur Turrell (@arthurturrell) September 15, 2026
The AI-as-Normal-Technology view of loss of control incidents
A middle ground between the cybersecurity and AI safety communities

The Alien Space of Science: Sampling Coherent but Cognitively...
Scientific discovery is constrained not only by what is true, but by what is cognitively available to the researchers currently exploring a field. Many directions are coherent in light of the...

CausalSmith · AI Causal Scientist
Working papers from CausalSmith, an AI causal scientist: econometric theory in which every theorem, assumption, and lemma is machine-verified in Lean 4.
Vasilis Syrgkanis on Twitter / X
Interesting times! Thanks to the amazing work of my student @JiyuanTan, we recently released a fully automated agentic pipeline that automates theoretical research in causal inference https://t.co/GclkTcHjjT.— Vasilis Syrgkanis (@syrgkanis) September 14, 2026
Sayash Kapoor on Twitter / X
What does it mean to pace the frontier? Over the last month, @random_walker and I have analyzed the loss-of-control incidents at AI companies to understand what technical and policy interventions can improve safety and what companies should do to pace the frontier. The result… pic.twitter.com/HB9dMZWbjy— Sayash Kapoor (@sayashk) September 14, 2026
Effort for Payment: A Tale of Two Markets
The standard model of labor is one in which individuals trade their time and energy in return for monetary rewards. Building on Fiske's relational theory (1992), we propose that there are two types of markets that determine relationships between effort and payment: monetary and social. We hypothesize that monetary markets are highly sensitive to the magnitude of compensation, whereas social markets are not. This perspective can shed light on the well-established observation that people sometimes expend more effort in exchange for no payment (a social market) than they expend when they receive low payment (a monetary market). Three experiments support these ideas. The experimental evidence also demonstrates that mixed markets (markets that include aspects of both social and monetary markets) more closely resemble monetary than social markets.

Build Agent Advocates, Not Platform Agents
Language model agents are poised to mediate how people navigate and act online. If the companies that already dominate internet search, communication, and commerce -- or the firms trying to unseat...

Lina Khan on Twitter / X
Law enforcers already have authority to charge companies and their CEOs for creating and releasing dangerous, unvetted, or defective products. We shouldn’t let discussions about new legal regimes distract from the fact that there’s no AI exemption from laws already on the books —…— Lina Khan (@linamkhan) September 13, 2026
The Mind in the Wheel
When the mind of its own in the wheel puts two and two together...
