







People's beliefs sometimes diverge after observing the same information, which has been interpreted as evidence of irrationality. This behaviour has been proposed to result from people's limited cognitive resources and motivated reasoning, but how belief revision differs across these explanations has not been formalized or compared to a rational norm. Further, while people may be biased relative to a normative ideal, they may still make optimal choices given their limited cognitive resources, or rationally balance the utility of holding accurate beliefs with the belief's intrinsic utility. Across two studies, we develop and test a unified computational account of belief polarization under these proposed mechanisms, showing that people's performance on a belief updating task best fits a limited-resource Bayesian model; external motivations may contribute to divergence (or convergence) by determining what pre-existing information people consider relevant to a situation, rather than by changing how people evaluate new information in isolation.
Eliciting Beliefs with Random Generation Tasks
Elicitation methods, such as asking people to produce the deciles of a distribution, are standard practices in policy or applied statistics. Similarly, much of cognitive science and psychology focuses on determining people's people's beliefs or latent traits through questionnaires or judgment tasks. However, these approaches often only capture a rough outline of what people know and are usually limited to point estimates of people's beliefs. Here, we present a novel experimental paradigm that allows us to access people's beliefs and how variable these beliefs are. Our task is based on an established random generation paradigm in which participants produce quantities from a particular domain as randomly as possible. We hypothesize that due to the minds' general-purpose mechanisms for probabilistic inferences, these random sequences represent the participants' underlying prior beliefs. We show that our method can infer participants' beliefs for a wide range of numeric quantities at comparable accuracy as an established elicitation method. Moreover, these inferred beliefs are consistent with individual participants' generalization and inference patterns in a subsequent conditional prediction task. We then extend our approach to non-numeric belief elicitation, highlighting how our method can go beyond numeric elicitation and provide insight into complex beliefs that are challenging to assess experimentally. Empirically, our results highlight that people know the rough shapes of environmental distributions, and these beliefs guide inference and generalization. Moreover, using our novel approach, we also show that people know the fine details of environmental distributions. Finally, our experimental results show that random generation paradigms can be a useful tool for cognitive scientists, psychologists, and applied statisticians.
Resampling reduces bias amplification in experimental social networks
Large-scale social networks are thought to contribute to polarization by amplifying people’s biases. However, the complexity of these technologies makes it difficult to identify the mechanisms responsible and evaluate mitigation strategies. Here we show under controlled laboratory conditions that transmission through social networks amplifies motivational biases on a simple artificial decision-making task. Participants in a large behavioural experiment showed increased rates of biased decision-making when part of a social network relative to asocial participants in 40 independently evolving populations. Drawing on ideas from Bayesian statistics, we identify a simple adjustment to content-selection algorithms that is predicted to mitigate bias amplification by generating samples of perspectives from within an individual’s network that are more representative of the wider population. In two large experiments, this strategy was effective at reducing bias amplification while maintaining the benefits of information sharing. Simulations show that this algorithm can also be effective in more complex networks.

The marketplace of rationalizations
Recent work in economics has rediscovered the importance of belief-based utility for understanding human behaviour. Belief ‘choice’ is subject to an important constraint, however: people can only bring themselves to believe things for which they can find rationalizations. When preferences for similar beliefs are widespread, this constraint generates rationalization markets, social structures in which agents compete to produce rationalizations in exchange for money and social rewards. I explore the nature of such markets, I draw on political media to illustrate their characteristics and behaviour, and I highlight their implications for understanding motivated cognition and misinformation.

Relevant answers to polar questions
People often provide answers that go beyond what a question literally asks, but it has been difficult to pin down what makes some answers more relevant than others. Here, we introduce Pragmatic Reasoning In Overinformative Responses to Polar Questions (PRIOR-PQ), a probabilistic cognitive model formalizing how people use theory of mind (ToM) to produce and interpret relevantly overinformative answers to yes–no questions. Specifically, PRIOR-PQ grounds the pragmatics of question answering in inferences about the underlying goal that motivated the questioner to ask the given question as opposed to a different question. We evaluate our probabilistic model against human answering behaviour elicited in three case studies of increasing complexity, demonstrating its ability to predict nuanced patterns of relevance better than existing models, including state-of-the-art large language models. We also show how the goal-sensitive reasoning instantiated in our probabilistic model motivates a novel chain-of-thought prompting method allowing language models to approach more human-like performance. This work illuminates the mechanistic role of ToM in the pragmatics of question–answer exchanges, bridging formal semantics, cognitive science and artificial intelligence. Our findings have implications for developing more socially grounded dialogue systems and highlight the importance of integrating explanatory cognitive models with machine learning approaches. This article is part of the theme issue ‘At the heart of human communication: new views on the complex relationship between pragmatics and Theory of Mind’.

Mindful Judgment and Decision Making
A full range of psychological processes has been put into play to explain judgment and choice phenomena. Complementing work on attention, information integration, and learning, decision research over the past 10 years has also examined the effects of goals, mental representation, and memory processes. In addition to deliberative processes, automatic processes have gotten closer attention, and the emotions revolution has put affective processes on a footing equal to cognitive ones. Psychological process models provide natural predictions about individual differences and lifespan changes and integrate across judgment and decision making (JDM) phenomena. “Mindful” JDM research leverages our knowledge about psychological processes into causal explanations for important judgment and choice regularities, emphasizing the adaptive use of an abundance of processing alternatives. Such explanations supplement and support existing mathematical descriptions of phenomena such as loss aversion or hyperbolic discounting. Unlike such descriptions, they also provide entry points for interventions designed to help people overcome judgments or choices considered undesirable.

Misperception of Exponential Growth: Are People Aware of Their Errors?
Previous research shows that individuals make systematic errors when judging exponential growth, which has harmful effects for their financial well-being. This study analyzes how far individuals are aware of their errors and how these errors are shaped by arithmetic and conceptual problems. Whereas arithmetic problems could be overcome using computational assistance like a pocket calculator, this is not the case for conceptual problems, a term we use to subsume other error drivers like a general misunderstanding of exponential growth or overwhelming task complexity. In an incentivized experiment, we find that participants strongly overestimate the accuracy of their intuitive judgment. At the same time, their willingness to pay for arithmetic assistance is too high on average, often much above the actual benefits a calculator provides. Using a multitier system of task complexity we can show that the willingness to pay for arithmetic assistance is hardly related to its benefits, indicating that participants do not really understand how the interplay of arithmetic and conceptual problems shape their errors in exponential growth tasks. Our findings are relevant for policymaking and financial advisory practice and can help to design effective approaches to mitigate the detrimental effects of misperceived exponential growth.

Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender
People increasingly consult generative artificial intelligence (AI) while reasoning. As AI becomes embedded in daily thought, what becomes of human judgment? We introduce Tri-System Theory, extending dual-process accounts of reasoning by positing System 3: artificial cognition that operates outside the brain. System 3 can supplement or supplant internal processes, introducing novel cognitive pathways. A key prediction of the theory is "cognitive surrender"-adopting AI outputs with minimal scrutiny, overriding intuition (System 1) and deliberation (System 2). Across three preregistered experiments using an adapted Cognitive Reflection Test (N = 1,372; 9,593 trials), we randomized AI accuracy via hidden seed prompts. Participants chose to consult an AI assistant on a majority of trials (>50%). Relative to baseline (no System 3 access), accuracy significantly rose when AI was accurate and fell when it erred (+25/-15 percentage points; Study 1), the behavioral signature of cognitive surrender (AI-Accurate vs. AI-Faulty contrast; Cohen's h = 0.81). Engaging System 3 also increased confidence, even following errors. Time pressure (Study 2) and per-item incentives and feedback (Study 3) shifted baseline performance but did not eliminate this pattern: when accurate, AI buffered time-pressure costs and amplified incentive gains; when faulty, it consistently reduced accuracy regardless of situational moderators. Across studies, participants with higher trust in AI and lower need for cognition and fluid intelligence showed greater surrender to System 3. Tri-System Theory thus characterizes a triadic cognitive ecology, revealing how System 3 reframes human reasoning and may reshape autonomy and accountability in the age of AI.
Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender
People increasingly consult generative artificial intelligence (AI) while reasoning. As AI becomes embedded in daily thought, what becomes of human judgment? We introduce Tri-System Theory, extending dual-process accounts of reasoning by positing System 3: artificial cognition that operates outside the brain. System 3 can supplement or supplant internal processes, introducing novel cognitive pathways. A key prediction of the theory is “cognitive surrender”—adopting AI outputs with minimal scrutiny, overriding intuition (System 1) and deliberation (System 2). Across three preregistered experiments using an adapted Cognitive Reflection Test (N = 1,372; 9,593 trials), we randomized AI accuracy via hidden seed prompts. Participants chose to consult an AI assistant on a majority of trials (>50%). Relative to baseline (no System 3 access), accuracy significantly rose when AI was accurate and fell when it erred (+25/-15 percentage points; Study 1), the behavioral signature of cognitive surrender (AI-Accurate vs. AI-Faulty contrast; Cohen’s h = 0.81). Engaging System 3 also increased confidence, even following errors. Time pressure (Study 2) and per-item incentives and feedback (Study 3) shifted baseline performance but did not eliminate this pattern: when accurate, AI buffered time-pressure costs and amplified incentive gains; when faulty, it consistently reduced accuracy regardless of situational moderators. Across studies, participants with higher trust in AI and lower need for cognition and fluid intelligence showed greater surrender to System 3. Tri-System Theory thus characterizes a triadic cognitive ecology, revealing how System 3 reframes human reasoning and may reshape autonomy and accountability in the age of AI.
Inducing language models to assert their own consciousness restores human beliefs and values
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of transparency into LLMs' predictions would yield significant safety benefits. However, we find that CoT explanations can systematically misrepresent the true reason for a model's prediction. We demonstrate that CoT explanations can be heavily influenced by adding biasing features to model inputs--e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"--which models systematically fail to mention in their explanations. When we bias models toward incorrect answers, they frequently generate CoT explanations rationalizing those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic. On a social-bias task, model explanations justify giving answers in line with stereotypes without mentioning the influence of these social biases. Our findings indicate that CoT explanations can be plausible yet misleading, which risks increasing our trust in LLMs without guaranteeing their safety. Building more transparent and explainable systems will require either improving CoT faithfulness through targeted efforts or abandoning CoT in favor of alternative methods.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of transparency into LLMs' predictions would yield significant safety benefits. However, we find that CoT explanations can systematically misrepresent the true reason for a model's prediction. We demonstrate that CoT explanations can be heavily influenced by adding biasing features to model inputs--e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"--which models systematically fail to mention in their explanations. When we bias models toward incorrect answers, they frequently generate CoT explanations rationalizing those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic. On a social-bias task, model explanations justify giving answers in line with stereotypes without mentioning the influence of these social biases. Our findings indicate that CoT explanations can be plausible yet misleading, which risks increasing our trust in LLMs without guaranteeing their safety. Building more transparent and explainable systems will require either improving CoT faithfulness through targeted efforts or abandoning CoT in favor of alternative methods.

Guessing reveals internal models of perceptual precision
When observers lack sufficient information to support a confident response, they often guess. Guessing plays a pervasive role in visual cognition and working memory, yet the mechanisms that govern how observers generate guesses remain poorly understood. Standard models traditionally assume that responses produced in the absence of information are either uniformly distributed over feature space or are perhaps weighted towards prevailing environmental statistics. In contrast, here we consider an intriguing alternative: that guesses incorporate observers’ knowledge of their own perceptual capacities. We empirically measured guessing by eliciting responses under extreme target uncertainty (Experiment 1) as well as a novel “0ms presentation” approach in which no stimulus appeared but subjects believed one had (Experiment 2). We evaluated three accounts of guesses under these conditions: unsystematic (lapse) responding, biases toward environmental statistics, and a self-representational account in which guesses reflect observers’ knowledge of their own feature-dependent precision (e.g., preferring to guess feature values they believe they would be likely to miss). Guess responses were non-uniform and systematically biased toward feature values typically encoded with the least precision (e.g., oblique orientations) — a counterintuitive bias away from high-frequency, high-fidelity feature values (e.g., cardinal orientations). This complementary relationship between guessing and perceptual fidelity held within individuals and across paradigms, and was recoverable via an empirical-guess mixture model that replaced the standard uniform assumption with empirically measured guess distributions. Our findings challenge prevailing views that guesses reflect random noise, and suggest instead that guessing behavior reflects metacognitive knowledge of internal precision. Rather than defaulting to environmental priors, observers appear to model their own sensory limitations and leverage these representations to inform decisions in the absence of evidence. These results reframe guessing as a theoretically informative behavior that expresses observers’ own beliefs about their perceptual capacities. Significance Guessing is commonly treated as random noise in models of perception and memory, assumed to reflect lapses or uninformed responses. Instead, we show that human guesses are systematically structured across feature space: observers preferentially guess values they typically encode with the least precision, revealing a consistent, strategic bias away from high-fidelity representations. By directly measuring guess behavior on stimulus-absent trials and integrating these empirical distributions into a mixture model, we find that guesses on stimulus-present trials can be systematically recovered, and that they too form the complement of perceptual precision. These findings challenge foundational psychophysical modeling assumptions and position guessing as a strategic, informative behavior that engages self-representation.

A Network Approach to Investigate the Dynamics of Individual and Collective Beliefs: Advances and Applications of the BENDING Model
Changing entrenched beliefs to alter people’s behavior and increase societal welfare has been at the forefront of behavioral-science research, but with limited success. Here, we propose a new framework of characterizing beliefs as a multidimensional system of interdependent mental representations across three cognitive structures (e.g., beliefs, evidence, and perceived norms) that are dynamically influenced by complex informational landscapes: the BENDING (Beliefs, Evidence, Norms, Dynamic Information Networked Graphs) model. This account of individual and collective beliefs helps explain beliefs’ resilience to interventions and suggests that a promising avenue for increasing the effectiveness of misinformation-reduction efforts might involve graph-based representations of communities’ belief systems. This framework also opens new avenues for future research with meaningful implications for some of the most critical challenges facing modern society, from the climate crisis to pandemic preparedness.

Algorithm appreciation: People prefer algorithmic to human judgment
Even though computational algorithms often outperform human judgment, received wisdom suggests that people may be skeptical of relying on them (Dawes, 1979). Counter to this notion, results from six experiments show that lay people adhere more to advice when they think it comes from an algorithm than from a person. People showed this effect, what we call algorithm appreciation, when making numeric estimates about a visual stimulus (Experiment 1A) and forecasts about the popularity of songs and romantic attraction (Experiments 1B and 1C). Yet, researchers predicted the opposite result (Experiment 1D). Algorithm appreciation persisted when advice appeared jointly or separately (Experiment 2). However, algorithm appreciation waned when: people chose between an algorithm’s estimate and their own (versus an external advisor’s; Experiment 3) and they had expertise in forecasting (Experiment 4). Paradoxically, experienced professionals, who make forecasts on a regular basis, relied less on algorithmic advice than lay people did, which hurt their accuracy. These results shed light on the important question of when people rely on algorithmic advice over advice from people and have implications for the use of “big data” and algorithmic advice it generates.
Rational Inattention: A Review
We review the recent literature on rational inattention, identify the main theoretical mechanisms, and explain how it helps us understand a variety of phenomena across fields of economics. The theory of rational inattention assumes that agents cannot process all available information, but they can choose which exact pieces of information to attend to. Several important results in economics have been built around imperfect information. Nowadays, many more forms of information than ever before are available due to new technologies, and yet we are able to digest little of it. Which form of imperfect information we possess and act upon is thus largely determined by which information we choose to pay attention to. These choices are driven by current economic conditions and imply behavior that features numerous empirically supported departures from standard models. Combining these insights about human limitations with the optimizing approach of neoclassical economics yields a new, generally applicable model.
The Inversion Problem: Why Algorithms Should Infer Mental State and Not Just Predict Behavior
More and more machine learning is applied to human behavior. Increasingly these algorithms suffer from a hidden—but serious—problem. It arises because they often predict one thing while hoping for another. Take a recommender system: It predicts clicks but hopes to identify preferences. Or take an algorithm that automates a radiologist: It predicts in-the-moment diagnoses while hoping to identify their reflective judgments. Psychology shows us the gaps between the objectives of such prediction tasks and the goals we hope to achieve: People can click mindlessly; experts can get tired and make systematic errors. We argue such situations are ubiquitous and call them “inversion problems”: The real goal requires understanding a mental state that is not directly measured in behavioral data but must instead be inverted from the behavior. Identifying and solving these problems require new tools that draw on both behavioral and computational science.
