







Our most up-to-date measurements of the time horizons for public frontier language models.
Large language models outperform humans at estimating society's everyday norms—but hybrids are even better
As AI assistants and social robots enter human environments, their ability to navigate context-dependent social norms is essential to avoid harm and ensure successful collaboration. We evaluate six large language models (LLMs) on their ability to estimate American social norms across 555 everyday scenarios (measured in prior work) and compare these to estimates from 320 humans. LLMs achieve remarkably high accuracy, clearly outperforming the average human. However, the errors LLMs make are systematic; they are similar across runs of the same LLM and even across different LLMs. As a consequence of this homogeneity, aggregating estimates of LLMs produces little improvement. Individual humans make much worse estimates, often defaulting to extreme right-or-wrong judgments even when asked to estimate population averages, but their errors are idiosyncratic and, consequently, aggregating their estimates yields dramatic improvement through wisdom-of-crowds effects. As humans make different errors than LLMs, hybrid ensembles combining both substantially outperform either alone.
Pacing the Frontier
A statement from over 1000 employees of frontier AI companies

Open-world evaluations for measuring frontier AI capabilities
Introducing CRUX, a new project for evaluating AI on long, messy tasks

The Philosophy of Language Models
ABSTRACT The success of large language models (LLMs) across many domains of AI research has generated intense debate. Some attribute their impressive performance on complex tasks to human‐like linguistic and cognitive capacities, whereas others ascribe it to shallow pattern matching. These disputes stem from deep‐seated philosophical disagreements about the nature of language and cognition. We provide an opinionated survey of these disagreements across core topics in the philosophy of mind and language, including syntactic competence, compositionality, linguistic meaning, representation, attitudes, reasoning, agency, and consciousness. We contend that progress on these issues requires not only clarity about background philosophical commitments but also, in many cases, close engagement with emerging empirical evidence.

Are AI Models on the Autism Spectrum? Exploring the Parallels
Large language models are celebrated for their ability to process information quickly and generate human-like responses. However, they…

Models overview
Claude is a family of state-of-the-art large language models developed by Anthropic. This guide introduces the available models and compares their performance.
How Large Language Models Actually Work
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
Recent generations of frontier language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes…

Stanford CS336 Language Modeling from Scratch I 2025
Extracting Training Data from Large Language Models
Nicholas Carlini, Google; Florian Tramèr, Stanford University; Eric Wallace, UC Berkeley; Matthew Jagielski, Northeastern University; Ariel Herbert-Voss, OpenAI and Harvard University; Katherine Lee and Adam Roberts, Google; Tom Brown, OpenAI; Dawn Song, UC Berkeley; Úlfar Erlingsson, Apple; Alina Oprea, Northeastern University; Colin Raffel, Google
Mixture-of-Experts (MoE): The Birth and Rise of Conditional Computation
What two decades of research taught us about sparse language models...

Recursive Language Models: the paradigm of 2026
How we plan to manage extremely long contexts
.png?v=c8c07d4bf43b)
Slop-Machine Future
The arc of large language models is mediocre, and it bends toward “target procurement”.

AI’s Memorization Crisis
Large language models don’t “learn”—they copy. And that could change everything for the tech industry.