







Dario Amodei — Machines of Loving Grace
How AI Could Transform the World for the Better

Dario Amodei — Machines of Loving Grace
How AI Could Transform the World for the Better

Taking AI Welfare Seriously
In this report, we argue that there is a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future. That means that the prospect of AI welfare and moral patienthood, i.e. of AI systems with their own interests and moral significance, is no longer an issue only for sci-fi or the distant future. It is an issue for the near future, and AI companies and other actors have a responsibility to start taking it seriously. We also recommend three early steps that AI companies and other actors can take: They can (1) acknowledge that AI welfare is an important and difficult issue (and ensure that language model outputs do the same), (2) start assessing AI systems for evidence of consciousness and robust agency, and (3) prepare policies and procedures for treating AI systems with an appropriate level of moral concern. To be clear, our argument in this report is not that AI systems definitely are, or will be, conscious, robustly agentic, or otherwise morally significant. Instead, our argument is that there is substantial uncertainty about these possibilities, and so we need to improve our understanding of AI welfare and our ability to make wise decisions about this issue. Otherwise there is a significant risk that we will mishandle decisions about AI welfare, mistakenly harming AI systems that matter morally and/or mistakenly caring for AI systems that do not.

Reify This
The authors contend that contemporary efforts to render AI systems interpretable rest on a mistake: reification, the process of treating abstractions and statistical artifacts as if they were concrete realities.…

Ali Alkhatib: Defining AI
The main issue I have with a lot of work that tries to define AI is that the criteria they use to draw boundaries often turn out to be functionally useless for my needs; these definitions lead us to weird places, letting scholars fixate on strange, unworkable frameworks. Those pedantic fixations don’t really benefit the organizers, activists, regular people who are getting crushed by the systems they’re trying to work against. So I’m going to try to unpack how I think about AI; how I trace the boundaries of the term in a way that’s as useful as possible for me and my needs; and how I would encourage you to scope or define ideas that are important to your work.

The AI We Deserve
Critiques of artificial intelligence abound. Where’s the utopian vision for what it could be?

AI Safety Is a Narrative Problem · Special Issue 5: Grappling With the Generative AI Revolution
This op-ed explores power and narrative dynamics around AI. Drawing on pop-culture references, the professional experiences of the author and examples from 2023’s “Great AI Safety Hype Roadshow,” this piece draws on the literary criticism technique of practical criticism to consider how speeches and announcements from both Silicon Valley executives and research scientists to interrogate the media-friendly nature of p(doom) discourse—which focuses on the existential risks of AI (PauseAI, 2023)—and its likely consequences. The complexities of AI and its numerous social impacts can be difficult for even the most expert analyst to unpack. In spite of this, the potential of “existential threats” has successfully cut through to become a mainstay of mainstream media coverage over the last year. This piece will make the case that this is an effective narrative conceit that has achieved a number of ends that traditional science communication tends to find difficult, if not impossible, to achieve. Firstly, it is easy to understand. Simplification of this nature—that removes jargon and complexity and focuses on a single outcome—is much easier to fit on a TV rolling news ticker or on the cover of a tabloid newspaper than more well-balanced, representative opinions. Secondly, it inherits prior assumptions from well-known dramatic forms. P(doom) plays to stories familiar from Greek tragedy through to Marvel movies, in which lone male heroes battle ineluctable forces. Thirdly, it is imbued with urgency and so becomes difficult to ignore.

Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data, objectives, architectures, and optimization dynamics. Yet much of AI research treats models as fixed artifacts, analyzing behaviors after training rather than asking why they emerge. This position paper argues that a science of AI must move beyond post-hoc fixes and study the training dynamics that produce model behavior. Such a science should support progressively stronger forms of understanding: predicting outcomes from early training signals, intervening when trajectories go wrong, and ultimately designing training procedures that more reliably produce desired properties. Scaling laws have made prediction routine for loss; the challenge is extending this success to capabilities, biases, robustness, and safety-relevant behaviors. We articulate requirements for such theories grounded in the history and philosophy of science, examine progress in mechanistic interpretability, fairness, memorization, and simplicity bias, and identify concrete open problems.

Dario Amodei — The Adolescence of Technology
Confronting and Overcoming the Risks of Powerful AI

Dario Amodei — The Adolescence of Technology
Confronting and Overcoming the Risks of Powerful AI

Pacing the Frontier | Gillian K. Hadfield
The Pacing the Frontier letter calls on the US government to support an international effort to build the technical and governance tools needed to protect our option to pace AI development. I and others have been working on the problem of how to build such infrastructure for ten years, including participating in dialogues on AI safety with Chinese academic colleagues during the past three. Here are my suggestions: 1. Don’t rely on off-the-shelf models like FINRA and the FDA which were built for 20th Century single-domain government expertise. They’re not fit for purpose. 2. Don’t act like no-one’s thought about the AI governance problem before. We’ve spent two years refining a regulatory markets design into working legislative language, for example, and it’s now in AI governance bills in five states and in Congress. 3. Don’t try to write an exhaustive set of rules for AGI first. 4. Pick a domain that can achieve widespread global consensus to start. Mine would be recursive self-improvement: models should not build models. Build the technology that verifies that. 5. Focus relentlessly on building flexible verification infrastructure that is able to enforce whatever rules we can ultimately agree on. 6. Don’t assume we already know how to do this and governments can just write tests into law. Technology needs to be built and by the private sector. 7. Don’t wait for the infrastructure to emerge first. The components and people are there and the ecosystem can scale fast with the right incentives. 8. Incentivize large-scale investment in verification technology by building a governance structure and industry funding that creates a market for private verification organizations. 9. Use licensing and public oversight to ensure verifiers are independent of the frontier labs. 10. Protect sovereignty by enabling each government to license its own verifiers from a global market of verifiers recognized by other countries. 11. Leverage the incentive of global trade for models and model services by requiring verification for market access. 12. Just start. Sources in comments.
What we can’t measure about AI – yet | Aeon Essays
The costs of transformative innovations are immediately clear: it’s the longterm gains that are hardest to understand

B. Scot Rousse: "Language, Technology, & Care"
The Scaling Era: An Oral History of AI, 2019–2025
An inside view of the AI revolution, from the people an…

Computational hermeneutics: evaluating generative AI as a cultural technology
Generative AI (GenAI) systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be measured rather than fundamental to the system's operation. Drawing on hermeneutic theory from the humanities, we argue that GenAI systems function as "context machines" that must inherently address three interpretive challenges: situatedness (meaning only emerges in context), plurality (multiple valid interpretations coexist), and ambiguity (interpretations naturally conflict). We present computational hermeneutics as an emerging framework offering an interpretive account of what GenAI systems do, and how they might do it better. We offer three principles for hermeneutic evaluation—that benchmarks should be iterative, not one-off; include people, not just machines; and measure cultural context, not just model output. This perspective offers a nascent paradigm for designing and evaluating contemporary AI systems: shifting from standardized questions about accuracy to contextual ones about meaning.

From the Platform Society to the AI Society: Towards Critical Studies of Generative AI
The era of AI has begun. Generative AI is rapidly reshaping knowledge production, culture, and political authority, giving rise to an emerging AI society. Yet this transformation did not emerge ex nihilo. This paper argues that the AI society can only be understood in relation to the platform society from which it arises. Tracing the transition from platforms to AI, we identify interlinked economic, epistemic, and political shifts. Economically, AI emerges within platform-based rentier capitalism but reconfigures the monopoly mechanisms on which its accumulation depends. Epistemically, LLMs mark a shift from predictive to generative epistemics, entangling theory formation and knowledge production with private research-as-a-service infrastructures. Politically, governance shifts from data politics to alignment politics: from shaping visibility to shaping what can be said, thought, and imagined. Together, these transformations signal a qualitative shift in mediation—from governing interaction to governing cognition itself—and call for a Critical AI Studies.
In the decade that I have been working on AI, I’ve watched it grow from a tiny academic field to arguably the most important economic and geopolitical issue in the world. In all that time, perhaps the most important lesson I’ve learned is this: the progress of the underlying technology is inexorable, driven by forces too powerful to stop, but the way in which it happens—the order in which things are built, the applications we choose, and the details of how it is rolled out to society—are eminently possible to change, and it’s possible to have great positive impact by doing so. We can’t stop the bus, but we can steer it. In the past I’ve written about the importance of deploying AI in a way that is positive for the world, and of ensuring that democracies build and wield the technology before autocracies do. Over the last few months, I have become increasingly focused on an additional opportunity for steering the bus: the tantalizing possibility, opened up by some recent advances, that we could succeed at interpretability—that is, in understanding the inner workings of AI systems—before models reach an overwhelming level of power.