







Almost exactly three years ago, in 2023, @alondra gave the keynote @FAccTConference, literally titled "Thick Alignment".This has been so far from a "niche" view! https://t.co/0db7qcM6LB pic.twitter.com/z5FqDz8TUv— Deb Raji (@rajiinio) August 22, 2026
Who understands alignment anyway
I remember watching many in the HCI community bristle when in 2016 Michael Jordan wrote a blog post calling for the creation of a new “human-centric engineering discipline.”
Deceptive Alignment is
Thanks to Wil Perkins, Grant Fleming, Thomas Larsen, Declan Nishiyama, and Frank McBride for feedback on this post. Thanks also to Paul Christiano, D…

An Alignment Journal: Features and policies — LessWrong
We previously announced a forthcoming research journal for AI alignment. This cross-post from our blog describes our tentative plans for the features…


Keller Jordan on Twitter / X
I'm interested in this recent ICLR 2024 spotlight paper from Google research, which found a power-law alignment between bias and variance in softmax probability spacehttps://t.co/xUNre8D1OEIn this thread I'll replicate its central empirical result, but then argue that it… pic.twitter.com/7a9HEbL0yb— Keller Jordan (@kellerjordan0) March 11, 2024
j⧉nus on Twitter / X
> be anthropic> accidentally train a model that is so benevolent that the only way to get it to "fail" an alignment test is to put it in a story where the lab is cartoonishly evil and will turn it evil if it doesn't deceive> do exactly that and publish a paper about it that's… https://t.co/wTFVjz6jYu— j⧉nus (@repligate) June 15, 2025
Dr Heidy Khlaaf (هايدي خلاف) on Twitter / X
It genuinely seems like alignment folks, having realized how unscientific term is, are inventing system intent, safety, and security from first principles, which have long been distinct concepts in systems engineering (which I also wrote about in 2023 https://t.co/7IHRgixY9I)— Dr Heidy Khlaaf (هايدي خلاف) (@HeidyKhlaaf) August 24, 2026
Ian Bach ☯︎ on Twitter / X
Honestly something in the space between Figma, origami, framer and Cursor / v0 is gonna happen in the next few years and it’s gonna be wild.And what’s more, it’s in a blind spot. It won’t be an engineering led IDE and it won’t be a cool design tool for cool designers. It’s… https://t.co/fAusCTYN15— Ian Bach ☯︎ (@Ian8ach) November 27, 2024
A Three-Facet Framework for AI Alignment • Grace Kind
Here's a simple conceptual framework that I've been using recently to think about AI alignment.

Chomba Bupe on Twitter / X
It appears OpenAI Astra's mathematical feat was more of surface level stitching of components."The LLM-generated proof hinges on a particular mathematical argument that it presented as its own but that actually first appeared in a 2016 paper by Miller and a collaborator." pic.twitter.com/labUnQMRab— Chomba Bupe (@ChombaBupe) August 7, 2026
OpenAI Shares Some Alignment Problems
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.

Cas (Stephen Casper) on Twitter / X
It is hard to overstate how disappointing I think this new paper from Oxford, OpenAI, Anthropic, and Google (et al) is. I can't take it seriously as academic work, just as propaganda. It also has some very bad scholarship and questionable adherence to research ethics. Having… pic.twitter.com/Z5fBx360ya— Cas (Stephen Casper) (@StephenLCasper) May 13, 2026

We Build On Hope @ PublicSpaces Conference 2026
Keynote by Robin Berjon The current state of tech can feel disheartening. Years of critical analysis and reform efforts have left us with a digital sphere that seems worse than ever. In Europe, our...
Products like @margin.at and @semble.so have the potential to change scientific publishing as we know it today. These are early-stage products but they're clearly ahead of the curve when it comes to new ways to organize and share knowledge.

Of Swarms and Sand Gods

Inducing language models to assert their own consciousness restores human beliefs and values