







Abstract Determining the knowledge that guides human judgments is fundamental to understanding how people reason, make decisions, and form predictions. We use an experimental procedure called ‘‘iterated learning,’’ in which the responses that people give on one trial are used to generate the data they see on the next, to pinpoint the knowledge that informs people's predictions about everyday events (e.g., predicting the total box office gross of a movie from its current take). In particular, we use this method to discriminate between two models of human judgments: a simple Bayesian model ( Griffiths & Tenenbaum, 2006 ) and a recently proposed alternative model that assumes people store only a few instances of each type of event in memory (Min K ; Mozer, Pashler, & Homaei, 2008 ). Although testing these models using standard experimental procedures is difficult due to differences in the number of free parameters and the need to make assumptions about the knowledge of individual learners, we show that the two models make very different predictions about the outcome of iterated learning. The results of an experiment using this methodology provide a rich picture of how much people know about the distributions of everyday quantities, and they are inconsistent with the predictions of the Min K model. The results suggest that accurate predictions about everyday events reflect relatively sophisticated knowledge on the part of individuals.
Eliciting Beliefs with Random Generation Tasks
Elicitation methods, such as asking people to produce the deciles of a distribution, are standard practices in policy or applied statistics. Similarly, much of cognitive science and psychology focuses on determining people's people's beliefs or latent traits through questionnaires or judgment tasks. However, these approaches often only capture a rough outline of what people know and are usually limited to point estimates of people's beliefs. Here, we present a novel experimental paradigm that allows us to access people's beliefs and how variable these beliefs are. Our task is based on an established random generation paradigm in which participants produce quantities from a particular domain as randomly as possible. We hypothesize that due to the minds' general-purpose mechanisms for probabilistic inferences, these random sequences represent the participants' underlying prior beliefs. We show that our method can infer participants' beliefs for a wide range of numeric quantities at comparable accuracy as an established elicitation method. Moreover, these inferred beliefs are consistent with individual participants' generalization and inference patterns in a subsequent conditional prediction task. We then extend our approach to non-numeric belief elicitation, highlighting how our method can go beyond numeric elicitation and provide insight into complex beliefs that are challenging to assess experimentally. Empirically, our results highlight that people know the rough shapes of environmental distributions, and these beliefs guide inference and generalization. Moreover, using our novel approach, we also show that people know the fine details of environmental distributions. Finally, our experimental results show that random generation paradigms can be a useful tool for cognitive scientists, psychologists, and applied statisticians.
Iterated learning: Intergenerational knowledge transmission reveals inductive biases
Cultural transmission of information plays a central role in shaping human knowledge. Some of the most complex knowledge that people acquire, such as languages or cultural norms, can only be learned from other people, who themselves learned from previous generations. The prevalence of this process of “iterated learning” as a mode of cultural transmission raises the question of how it affects the information being transmitted. Analyses of iterated learning utilizing the assumption that the learners are Bayesian agents predict that this process should converge to an equilibrium that reflects the inductive biases of the learners. An experiment in iterated function learning with human participants confirmed this prediction, providing insight into the consequences of intergenerational knowledge transmission and a method for discovering the inductive biases that guide human inferences.
Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
Despite widespread success in language understanding and generation, large language models (LLMs) exhibit unclear and often inconsistent behavior when faced with tasks that require probabilistic reasoning. In this work, we present the first comprehensive study of the reasoning capabilities of LLMs over explicit discrete probability distributions. Given observations from a probability distribution, we evaluate models on three carefully designed tasks, mode identification, maximum likelihood estimation, and sample generation, by prompting them to provide responses to queries about either the joint distribution or its conditionals. These tasks thus probe a range of probabilistic skills, including frequency analysis, marginalization, and generative behavior. Through comprehensive empirical evaluations, we demonstrate that there exists a clear performance gap between smaller and larger models, with the latter demonstrating stronger inference and surprising capabilities in sample generation. Furthermore, our investigations reveal notable limitations, including sensitivity to variations in the notation utilized to represent probabilistic outcomes and performance degradation of over 60% as context length increases. Together, our results provide a detailed understanding of the probabilistic reasoning abilities of LLMs and identify key directions for future improvement.

Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
Despite widespread success in language understanding and generation, large language models (LLMs) exhibit unclear and often inconsistent behavior when faced with tasks that require probabilistic reasoning. In this work, we present the first comprehensive study of the reasoning capabilities of LLMs over explicit discrete probability distributions. Given observations from a probability distribution, we evaluate models on three carefully designed tasks, mode identification, maximum likelihood estimation, and sample generation, by prompting them to provide responses to queries about either the joint distribution or its conditionals. These tasks thus probe a range of probabilistic skills, including frequency analysis, marginalization, and generative behavior. Through comprehensive empirical evaluations, we demonstrate that there exists a clear performance gap between smaller and larger models, with the latter demonstrating stronger inference and surprising capabilities in sample generation. Furthermore, our investigations reveal notable limitations, including sensitivity to variations in the notation utilized to represent probabilistic outcomes and performance degradation of over 60% as context length increases. Together, our results provide a detailed understanding of the probabilistic reasoning abilities of LLMs and identify key directions for future improvement.

Language and Experience: A Computational Model of Social Learning in Complex Tasks
The ability to combine linguistic guidance from others with direct experience is central to human development, enabling safe and rapid learning in new environments. How do people integrate these two sources of knowledge, and how might AI systems? We present a computational framework that models social learning as joint probabilistic inference over structured, executable world models given sensorimotor and linguistic data. We make this possible by turning a pretrained language model into a probabilistic model of how humans share advice conditioned on their beliefs, allowing our agents both to generate advice for others and to interpret linguistic input as evidence during Bayesian inference. Using behavioral experiments and simulations across 10 video games, we show how linguistic guidance can shape exploration and accelerate learning by reducing risky interactions and speeding up key discoveries in both humans and models. We further explore how knowledge can accumulate across generations through iterated learning experiments and demonstrate successful knowledge transfer between humans and models -- revealing how structured, language-compatible representations might enable human-machine collaborative learning.

From Predictive Algorithms to Automatic Generation of Anomalies
Machine learning algorithms can find predictive signals that researchers fail to notice; yet they are notoriously hard-to-interpret. How can we extract theoretical insights from these black boxes? History provides a clue. Facing a similar problem – how to extract theoretical insights from their intuitions – researchers often turned to “anomalies:” constructed examples that highlight flaws in an existing theory and spur the development of new ones. Canonical examples include the Allais paradox and the Kahneman-Tversky choice experiments for expected utility theory. We suggest anomalies can extract theoretical insights from black box predictive algorithms. We develop procedures to automatically generate anomalies for an existing theory when given a predictive algorithm. We cast anomaly generation as an adversarial game between a theory and a falsifier, the solutions to which are anomalies: instances where the black box algorithm predicts - were we to collect data - we would likely observe violations of the theory. As an illustration, we generate anomalies for expected utility theory using a large, publicly available dataset on real lottery choices. Based on an estimated neural network that predicts lottery choices, our procedures recover known anomalies and discover new ones for expected utility theory. In incentivized experiments, subjects violate expected utility theory on these algorithmically generated anomalies; moreover, the violation rates are similar to observed rates for the Allais paradox and Common ratio effect.

Prior Modeling
In Bayesian inference the prior model provides a valuable opportunity to incorporate domain expertise into our inferences. Unfortunately this opportunity often becomes a contentious issue in many fields, and this potential value is lost in the debate. In this case study I will discuss the challenges of building prior models that capture meaningful domain expertise and some practical strategies for ameliorating those challenges as much as possible.
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.

Against theory-motivated experimentation: Can random experimental choice lead to better theories?
Scientists must choose which among many experiments to perform. We study the epistemic success of experimental choice strategies proposed by philosophers of science or executed by scientists themselves. We develop a multi-agent model of the scientific process that jointly formalizes its core aspects: active experimentation, theorizing, and social learning. We find that agents who choose new experiments at random develop the most informative and predictive theories of the world. The agents aiming to confirm, falsify theories, or resolve theoretical disagreements end up with an illusion of epistemic success: they develop promising accounts for the data they collected, while misrepresenting the ground truth that they intended to learn about. Agents experimenting in these theory-motivated ways acquire less diverse or less representative samples from the ground truth that also turn out to be easier to account for. Random data collection, on the other hand, combines virtues of diverse and representative sampling from a target scientific domain which enables cumulative development of the successful theoretical accounts of it. We suggest that randomization, already a gold standard within experiments, is also beneficial at the level of experiments themselves.

Do People Ask Good Questions?
People ask questions in order to efficiently learn about the world. But do people ask good questions? In this work, we designed an intuitive, game-based task that allowed people to ask natural language questions to resolve their uncertainty. Question quality was measured through Bayesian ideal observer models that considered large spaces of possible game states. During free-form question generation, participants asked a creative variety of useful and goal-directed questions, yet they rarely asked the best questions as identified by the Bayesian ideal observers (Experiment 1). In subsequent experiments, participants strongly preferred the best questions when evaluating questions that they did not generate themselves (Experiments 2 and 3). On one hand, our results show that people can accurately evaluate question quality, even when the set of questions is diverse and an ideal observer analysis has large computational requirements. On the other hand, people have a limited ability to synthesize maximally informative questions from scratch, suggesting a bottleneck in the question asking process.

Exploring Psychology in the Field: Steps and Examples From the Used‐Car Market
Abstract The growing availability of large datasets in a variety of domains presents an opportunity for researchers to use field data to better understand psychological concepts. I discuss, from an empirical economics point of view, steps for how to study cognition in large datasets. I use two recent papers that explore psychology in the used‐car market as motivating examples. These examples help illustrate the potential importance of big data as a way to explore human psychology and cognition. , The growing availability of large datasets in a variety of domains presents an opportunity for researchers to use field data to better understand psychological concepts. I discuss from an empirical economics point of view, steps for how to study cognition in large datasets and illustrate these steps with recent empirical papers.

The Inversion Problem: Why Algorithms Should Infer Mental State and Not Just Predict Behavior
More and more machine learning is applied to human behavior. Increasingly these algorithms suffer from a hidden—but serious—problem. It arises because they often predict one thing while hoping for another. Take a recommender system: It predicts clicks but hopes to identify preferences. Or take an algorithm that automates a radiologist: It predicts in-the-moment diagnoses while hoping to identify their reflective judgments. Psychology shows us the gaps between the objectives of such prediction tasks and the goals we hope to achieve: People can click mindlessly; experts can get tired and make systematic errors. We argue such situations are ubiquitous and call them “inversion problems”: The real goal requires understanding a mental state that is not directly measured in behavioral data but must instead be inverted from the behavior. Identifying and solving these problems require new tools that draw on both behavioral and computational science.

Guessing reveals internal models of perceptual precision
When observers lack sufficient information to support a confident response, they often guess. Guessing plays a pervasive role in visual cognition and working memory, yet the mechanisms that govern how observers generate guesses remain poorly understood. Standard models traditionally assume that responses produced in the absence of information are either uniformly distributed over feature space or are perhaps weighted towards prevailing environmental statistics. In contrast, here we consider an intriguing alternative: that guesses incorporate observers’ knowledge of their own perceptual capacities. We empirically measured guessing by eliciting responses under extreme target uncertainty (Experiment 1) as well as a novel “0ms presentation” approach in which no stimulus appeared but subjects believed one had (Experiment 2). We evaluated three accounts of guesses under these conditions: unsystematic (lapse) responding, biases toward environmental statistics, and a self-representational account in which guesses reflect observers’ knowledge of their own feature-dependent precision (e.g., preferring to guess feature values they believe they would be likely to miss). Guess responses were non-uniform and systematically biased toward feature values typically encoded with the least precision (e.g., oblique orientations) — a counterintuitive bias away from high-frequency, high-fidelity feature values (e.g., cardinal orientations). This complementary relationship between guessing and perceptual fidelity held within individuals and across paradigms, and was recoverable via an empirical-guess mixture model that replaced the standard uniform assumption with empirically measured guess distributions. Our findings challenge prevailing views that guesses reflect random noise, and suggest instead that guessing behavior reflects metacognitive knowledge of internal precision. Rather than defaulting to environmental priors, observers appear to model their own sensory limitations and leverage these representations to inform decisions in the absence of evidence. These results reframe guessing as a theoretically informative behavior that expresses observers’ own beliefs about their perceptual capacities. Significance Guessing is commonly treated as random noise in models of perception and memory, assumed to reflect lapses or uninformed responses. Instead, we show that human guesses are systematically structured across feature space: observers preferentially guess values they typically encode with the least precision, revealing a consistent, strategic bias away from high-fidelity representations. By directly measuring guess behavior on stimulus-absent trials and integrating these empirical distributions into a mixture model, we find that guesses on stimulus-present trials can be systematically recovered, and that they too form the complement of perceptual precision. These findings challenge foundational psychophysical modeling assumptions and position guessing as a strategic, informative behavior that engages self-representation.

Talking with strangers is surprisingly informative
A meaningful amount of people’s knowledge comes from their conversations with others. The amount people expect to learn predicts their interest in having a conversation (pretests 1 and 2), suggesting that the presumed information value of conversations guides decisions of whom to talk with. The results of seven experiments, however, suggest that people may systematically underestimate the informational benefit of conversation, creating a barrier to talking with—and hence learning from—others in daily life. Participants who were asked to talk with another person expected to learn significantly less from the conversation than they actually reported learning afterward, regardless of whether they had conversation prompts and whether they had the goal to learn (experiments 1 and 2). Undervaluing conversation does not stem from having systematically poor opinions of how much others know (experiment 3) but is instead related to the inherent uncertainty involved in conversation itself. Consequently, people underestimate learning to a lesser extent when uncertainty is reduced, as in a nonsocial context (surfing the web, experiment 4); when talking to an acquainted conversation partner (experiment 5); and after knowing the content of the conversation (experiment 6). Underestimating learning in conversation is distinct from underestimating other positive qualities in conversation, such as enjoyment (experiment 7). Misunderstanding how much can be learned in conversation could keep people from learning from others in daily life.

Benchmarking World-Model Learning with Environment-Level Queries
World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, such as next-frame prediction or task return, and (ii) do not test whether a learned model supports diverse queries about the environment. In contrast, humans build $\textit{general-purpose}$ models that can answer many different questions about an environment$\unicode{x2014}$including questions that require understanding global structure and counterfactual consequences. We propose $\textit{WorldTest}$: a protocol for evaluating whether agents learn models that support multiple $\textit{environment-level queries}\unicode{x2014}$questions whose answers depend on properties of the full environment, not just observed trajectories. Individually, these queries can target properties (e.g., reachability or the effects of interventions) that no single rollout distribution determines. Collectively, they assess model generality across query types. We instantiate WorldTest as $\textit{AutumnBench}$, a benchmark of 43 interactive grid-world environments and 129 tasks across three query families for both humans and learning agents. Experiments with 517 human participants and five frontier models show that humans substantially outperform these models, a gap we attribute to differences in exploration and belief updating. AutumnBench provides a framework for evaluating world-model learning in grid-world environments with environment-level queries, and WorldTest provides a template for extending such evaluations to richer domains.

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory benchmarks primarily target declarative memory, specifically semantic and episodic types, where all information is explicitly presented in dialogues. In contrast, real-world actions are also governed by non-declarative memory, including habitual and procedural types, and need to be inferred from diverse digital traces. To bridge this gap, we introduce Lifebench, which features densely connected, long-horizon event simulation. It pushes AI agents beyond simple recall, requiring the integration of declarative and non-declarative memory reasoning across diverse and temporally extended contexts. Building such a benchmark presents two key challenges: ensuring data quality and scalability. We maintain data quality by employing real-world priors, including anonymized social surveys, map APIs, and holiday-integrated calendars, thus enforcing fidelity, diversity and behavioral rationality within the dataset. Towards scalability, we draw inspiration from cognitive science and structure events according to their partonomic hierarchy; enabling efficient parallel generation while maintaining global coherence. Performance results show that top-tier, state-of-the-art memory systems reach just 55.2\% accuracy, highlighting the inherent difficulty of long-horizon retrieval and multi-source integration within our proposed benchmark. The dataset and data synthesis code are available at https://github.com/1754955896/LifeBench.
