







Anthropic's latest interpretability research: a new microscope to understand Claude's internal mechanisms
A global workspace in language models
Interpretability research on Claude's internal thoughts.

Emotion concepts and their function in a large language model
Interpretability research from Anthropic on emotion concepts

Anthropic on Twitter / X
In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked how the values Claude expresses vary between Claude models and across languages.We analyzed 300K+ anonymized conversations to find out.https://t.co/PgxsMXipt5— Anthropic (@AnthropicAI) July 13, 2026
Models overview
Claude is a family of state-of-the-art large language models developed by Anthropic. This guide introduces the available models and compares their performance.
Understanding Understanding: A Pragmatic Framework Motivated by...
Motivated by the rapid ascent of Large Language Models (LLMs) and debates about the extent to which they possess human-level qualities, we propose a framework for testing whether any agent (be it...

Anthropic on Twitter / X
New Anthropic research: A global workspace in language models.Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.We found a strikingly similar divide inside Claude. pic.twitter.com/aLUPBifxth— Anthropic (@AnthropicAI) July 6, 2026
Mapping the Mind of a Large Language Model
We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model.


Large language models are cultural technologies. What might that mean?
Four different perspectives

On the Biology of a Large Language Model
We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology.

The assistant axis: situating and stabilizing the character of large language models
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

A Rational Analysis of the Effects of Sycophantic AI
People increasingly use large language models (LLMs) to explore ideas, gather information, and make sense of the world. In these interactions, they encounter agents that are overly agreeable. We...




