







KillBench: Discovering Hidden Biases of LLMs
1M+ experiments exposing bias in critical AI decision-making

AI Sycophancy and Decisions
We examine whether sycophantic AI advice distorts decisions. Our experiment involves 1,500 participants in 30 decision environments spanning core domains in eco
The Cosmos and the Model
Humboldt, the Romantics, and What AI Loses by Averaging

AI has supercharged scientists—but may have shrunk science
Analysis of 41 million papers finds that although AI expands individual impact, it narrows collective scientific exploration

Ethical AI Departures — Who Quit OpenAI, Google, Anthropic Over Safety Concerns
59 researchers and executives who quit or were fired from OpenAI, Google DeepMind, Anthropic, Meta, and xAI over AI safety concerns. Sourced departures, writings, and prediction tracking.
Algorithmic Bias · Open Encyclopedia of Cognitive Science
Algorithmic bias refers to prejudicial, discriminatory, unjust, inaccurate, or otherwise disparate performance or outcomes from algorithmic systems based on racial, gender, or other attributes of an individual or a group. The concept of algorithmic bias emerged at the intersection of computer science, artificial intelligence (AI) research, critical data studies, human–computer interaction, law, philosophy, and similar disciplines. Although problems and discrepancies at the model level denote the most commonly studied form of bias, the term algorithmic bias is also used as a shorthand to describe a multitude of problems and challenges at various steps of the AI pipeline from ideation, problem framing, training data curation and processing, model training and validation, and deployment as well as emergent issues that arise from interaction with the real world. Potential sources of bias, appropriate metrics to define, measure, and mitigate bias, and the utility and merit of technical approaches to bias mitigation are fiercely debated in the current AI landscape.

AI FOR EPISTEMICS & COORDINATION
Civilization and technology have radically improved the human condition. Nonetheless, the world sometimes goes in directions which essentially nobody would prefer — e.g., nuclear arms races, unexpected financial crashes, predatory marketing, or ubiquitous political misinformation.
AI Values Dashboard
How leading AI models value different people, companies, countries, and groups.
Who earned the score? - Sensemaker
This week's AI claims blurred models, systems, simulations and people. The evidence becomes clearer when the tested subject comes first.
Pacing the Frontier
A statement from over 1000 employees of frontier AI companies

The AI Model That Was Too Dangerous to Release: Meet Claude Mythos
Anthropic just built the most powerful AI in history. Then decided the world wasn’t ready for it. Here’s everything you need to know.
Why I don’t trust most human-AI interaction experimental research – Jason Collins blog
Behavioural economics, data science and artificial intelligence.
Large AI models are cultural and social technologies – Henry Farrell
Debates about artificial intelligence (AI) tend to revolve around whether large models are intelligent, autonomous agents. Some AI researchers and commentators speculate that we are on the cusp of creating agents with artificial general intelligence (AGI), a prospect anticipated with both elation and anxiety. There have also been extensive conversations about cultural and social consequences of large models, orbiting around two foci: immediate effects of these systems as they are currently used, and hypothetical futures when these systems turn into AGI agents perhaps even superintelligent AGI agents.

New research: how well do AI models actually follow their constitutions? 205 tenets from Anthropic's 30K-word soul doc. Adversarial multi-turn scenarios against 7 models. Claude: 15% → 2% violation rate in two generations. Training works. But the remaining failures tell a more important story.
Here’s a group I think it’s crucial for everyone to get to know: the AI successionists. They believe AI will be a "worthy successor." And they actually want it to replace humanity. They’re more influential than you might think! So I reported this for @vox.com 🧵 1/n