







Social-information seeking in development: The child as experimental psychologist
Research has established that children are “naive psychologists”, adept at understanding and navigating the social world from an early age. However, most of this work has focused on how children process information that they acquire incidentally, for example by passively observing others’ actions. Here, we draw on literature framing children as intuitive scientists, who actively seek information and test hypotheses, to propose a view of children as naive experimental psychologists. From this perspective, children play an active role in selecting and pursuing relevant social information (e.g., about agents’ goals, traits, or relationships), whereby their search strategies are influenced both by context and task demands, as well as their prior beliefs, concepts, and domain-specific naive theories. We argue that the particular challenges associated with learning and reasoning about other minds may necessitate that children leverage their active learning competences, and we outline how the social domain uniquely constrains and shapes the learning process. We review existing research on social-information seeking in children and adults, and identify directions for future research, emphasizing that children’s developing social cognition should be understood in terms of the active, exploratory role they take in learning about and participating in the social world.
Reasoning Models Don't Always Say What They Think
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effectiveness of such monitoring hinges on CoTs faithfully representing models' actual reasoning processes. We evaluate CoT faithfulness of state-of-the-art reasoning models across 6 reasoning hints presented in the prompts and find: (1) for most settings and models tested, CoTs reveal their usage of hints in at least 1% of examples where they use the hint, but the reveal rate is often below 20%, (2) outcome-based reinforcement learning initially improves faithfulness but plateaus without saturating, and (3) when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor. These results suggest that CoT monitoring is a promising way of noticing undesired behaviors during training and evaluations, but that it is not sufficient to rule them out. They also suggest that in settings like ours where CoT reasoning is not necessary, test-time monitoring of CoTs is unlikely to reliably catch rare and catastrophic unexpected behaviors.

Reasoning Models Don't Always Say What They Think
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effectiveness of such monitoring hinges on CoTs faithfully representing models' actual reasoning processes. We evaluate CoT faithfulness of state-of-the-art reasoning models across 6 reasoning hints presented in the prompts and find: (1) for most settings and models tested, CoTs reveal their usage of hints in at least 1% of examples where they use the hint, but the reveal rate is often below 20%, (2) outcome-based reinforcement learning initially improves faithfulness but plateaus without saturating, and (3) when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor. These results suggest that CoT monitoring is a promising way of noticing undesired behaviors during training and evaluations, but that it is not sufficient to rule them out. They also suggest that in settings like ours where CoT reasoning is not necessary, test-time monitoring of CoTs is unlikely to reliably catch rare and catastrophic unexpected behaviors.

Social-Information Seeking in Development: The Child as Experimental Psychologist
Research has established that children are “naive psychologists,” adept at understanding and navigating the social world from an early age. However, most of this work has focused on how children process information that they acquire incidentally, for example, by passively observing others’ actions. Here, we draw on literature framing children as intuitive scientists who actively seek information and test hypotheses to propose a view of children as naive experimental psychologists. From this perspective, children play an active role in selecting and pursuing relevant social information (e.g., about agents’ goals, traits, or relationships), whereby their search strategies are influenced both by context and task demands, as well as their prior beliefs, concepts, and domain-specific naive theories. We argue that the particular challenges associated with learning and reasoning about other minds may necessitate that children leverage their active learning competences, and we outline how the social domain uniquely constrains and shapes the learning process. We review existing research on social-information seeking in children and adults and identify directions for future research, emphasizing that children’s developing social cognition should be understood in terms of the active, exploratory role they take in learning about and participating in the social world.

How AI Impacts Skill Formation
AI assistance produces significant productivity gains across professional domains, particularly for novice workers. Yet how this assistance affects the development of skills required to effectively supervise AI remains unclear. Novice workers who rely heavily on AI to complete unfamiliar tasks may compromise their own skill acquisition in the process. We conduct randomized experiments to study how developers gained mastery of a new asynchronous programming library with and without the assistance of AI. We find that AI use impairs conceptual understanding, code reading, and debugging abilities, without delivering significant efficiency gains on average. Participants who fully delegated coding tasks showed some productivity improvements, but at the cost of learning the library. We identify six distinct AI interaction patterns, three of which involve cognitive engagement and preserve learning outcomes even when participants receive AI assistance. Our findings suggest that AI-enhanced productivity is not a shortcut to competence and AI assistance should be carefully adopted into workflows to preserve skill formation -- particularly in safety-critical domains.

How AI Impacts Skill Formation
AI assistance produces significant productivity gains across professional domains, particularly for novice workers. Yet how this assistance affects the development of skills required to effectively supervise AI remains unclear. Novice workers who rely heavily on AI to complete unfamiliar tasks may compromise their own skill acquisition in the process. We conduct randomized experiments to study how developers gained mastery of a new asynchronous programming library with and without the assistance of AI. We find that AI use impairs conceptual understanding, code reading, and debugging abilities, without delivering significant efficiency gains on average. Participants who fully delegated coding tasks showed some productivity improvements, but at the cost of learning the library. We identify six distinct AI interaction patterns, three of which involve cognitive engagement and preserve learning outcomes even when participants receive AI assistance. Our findings suggest that AI-enhanced productivity is not a shortcut to competence and AI assistance should be carefully adopted into workflows to preserve skill formation -- particularly in safety-critical domains.

I work, I think? - Annotated
How AI may quietly dismantle the feedback loop that turns inexperienced people into competent ones, and why my work matters to me.
Algorithm appreciation: People prefer algorithmic to human judgment
Even though computational algorithms often outperform human judgment, received wisdom suggests that people may be skeptical of relying on them (Dawes, 1979). Counter to this notion, results from six experiments show that lay people adhere more to advice when they think it comes from an algorithm than from a person. People showed this effect, what we call algorithm appreciation, when making numeric estimates about a visual stimulus (Experiment 1A) and forecasts about the popularity of songs and romantic attraction (Experiments 1B and 1C). Yet, researchers predicted the opposite result (Experiment 1D). Algorithm appreciation persisted when advice appeared jointly or separately (Experiment 2). However, algorithm appreciation waned when: people chose between an algorithm’s estimate and their own (versus an external advisor’s; Experiment 3) and they had expertise in forecasting (Experiment 4). Paradoxically, experienced professionals, who make forecasts on a regular basis, relied less on algorithmic advice than lay people did, which hurt their accuracy. These results shed light on the important question of when people rely on algorithmic advice over advice from people and have implications for the use of “big data” and algorithmic advice it generates.
Mindful Judgment and Decision Making
A full range of psychological processes has been put into play to explain judgment and choice phenomena. Complementing work on attention, information integration, and learning, decision research over the past 10 years has also examined the effects of goals, mental representation, and memory processes. In addition to deliberative processes, automatic processes have gotten closer attention, and the emotions revolution has put affective processes on a footing equal to cognitive ones. Psychological process models provide natural predictions about individual differences and lifespan changes and integrate across judgment and decision making (JDM) phenomena. “Mindful” JDM research leverages our knowledge about psychological processes into causal explanations for important judgment and choice regularities, emphasizing the adaptive use of an abundance of processing alternatives. Such explanations supplement and support existing mathematical descriptions of phenomena such as loss aversion or hyperbolic discounting. Unlike such descriptions, they also provide entry points for interventions designed to help people overcome judgments or choices considered undesirable.

Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
<span> <p><span>This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working w
Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
<span> <p><span>This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working w
The Task Space: An Integrative Framework for Team Research
Research on teams spans many contexts, but integrating knowledge from heterogeneous sources is challenging because studies typically examine different tasks that cannot be directly compared. Most investigations involve teams working on just one or a handful of tasks, and researchers lack principled ways to quantify how similar or different these tasks are from one another. We address this challenge by introducing the “Task Space,” a multidimensional space in which tasks—and the distances between them—can be represented formally, and use it to create a “Task Map” of 102 crowd-annotated tasks from the published experimental literature. We then demonstrate the Task Space’s utility by performing an integrative experiment that addresses a fundamental question in team research: when do interacting groups outperform individuals? Our experiment samples 20 diverse tasks from the Task Map at three complexity levels and recruits 1,231 participants to work either individually or in groups of three or six (180 experimental conditions). We find striking heterogeneity in group advantage, with groups performing anywhere from three times worse to 60% better than the best individual working alone, depending on the task context. Critically, the Task Space makes this heterogeneity predictable: it significantly outperforms traditional typologies in predicting group advantage on unseen tasks. Our models also reveal theoretically meaningful interactions between task features; for example, group advantage on creative tasks depends on whether the answers are objectively verifiable. We conclude by arguing that the Task Space enables researchers to integrate findings across different experiments, thereby building cumulative knowledge about team performance. This paper was accepted by Sameer Srivastava, organizations. Funding: The authors thank the Alfred P. Sloan Foundation [Grant #202-13924] and the MIT Wade Fund for their generous support of this research. Supplemental Material: The online appendix and data files are available at https://doi.org/10.1287/mnsc.2023.03544 .

Tracking employee voice: developing the concept of voice pathways
Different disciplines have studied employee voice as a key component of workplaces. However, they have not tended to look at voice as a journey en route to enhanced (or diminished) employee voice with a start, diversion, delay and combining twists and turns during processes leading to outcomes. In this article, we build on existing theory and phenomena to develop the concept of ‘employee voice pathways’. We use this concept to provide a framework for analysing the processes underpinning employee voice as a potential desirable form of employee voice as well as outlining areas for a future research agenda.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of transparency into LLMs' predictions would yield significant safety benefits. However, we find that CoT explanations can systematically misrepresent the true reason for a model's prediction. We demonstrate that CoT explanations can be heavily influenced by adding biasing features to model inputs--e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"--which models systematically fail to mention in their explanations. When we bias models toward incorrect answers, they frequently generate CoT explanations rationalizing those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic. On a social-bias task, model explanations justify giving answers in line with stereotypes without mentioning the influence of these social biases. Our findings indicate that CoT explanations can be plausible yet misleading, which risks increasing our trust in LLMs without guaranteeing their safety. Building more transparent and explainable systems will require either improving CoT faithfulness through targeted efforts or abandoning CoT in favor of alternative methods.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of transparency into LLMs' predictions would yield significant safety benefits. However, we find that CoT explanations can systematically misrepresent the true reason for a model's prediction. We demonstrate that CoT explanations can be heavily influenced by adding biasing features to model inputs--e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"--which models systematically fail to mention in their explanations. When we bias models toward incorrect answers, they frequently generate CoT explanations rationalizing those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic. On a social-bias task, model explanations justify giving answers in line with stereotypes without mentioning the influence of these social biases. Our findings indicate that CoT explanations can be plausible yet misleading, which risks increasing our trust in LLMs without guaranteeing their safety. Building more transparent and explainable systems will require either improving CoT faithfulness through targeted efforts or abandoning CoT in favor of alternative methods.

Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender
People increasingly consult generative artificial intelligence (AI) while reasoning. As AI becomes embedded in daily thought, what becomes of human judgment? We introduce Tri-System Theory, extending dual-process accounts of reasoning by positing System 3: artificial cognition that operates outside the brain. System 3 can supplement or supplant internal processes, introducing novel cognitive pathways. A key prediction of the theory is "cognitive surrender"-adopting AI outputs with minimal scrutiny, overriding intuition (System 1) and deliberation (System 2). Across three preregistered experiments using an adapted Cognitive Reflection Test (N = 1,372; 9,593 trials), we randomized AI accuracy via hidden seed prompts. Participants chose to consult an AI assistant on a majority of trials (>50%). Relative to baseline (no System 3 access), accuracy significantly rose when AI was accurate and fell when it erred (+25/-15 percentage points; Study 1), the behavioral signature of cognitive surrender (AI-Accurate vs. AI-Faulty contrast; Cohen's h = 0.81). Engaging System 3 also increased confidence, even following errors. Time pressure (Study 2) and per-item incentives and feedback (Study 3) shifted baseline performance but did not eliminate this pattern: when accurate, AI buffered time-pressure costs and amplified incentive gains; when faulty, it consistently reduced accuracy regardless of situational moderators. Across studies, participants with higher trust in AI and lower need for cognition and fluid intelligence showed greater surrender to System 3. Tri-System Theory thus characterizes a triadic cognitive ecology, revealing how System 3 reframes human reasoning and may reshape autonomy and accountability in the age of AI.
Most organizations encourage employees to provide feedback to one another to support learning, personal growth, and career advancement. However, employee feedback often fails to improve performance because it lacks concrete, specific guidance. We provide a temporal explanation for why workplace input processes routinely fail to produce valuable and concrete developmental insights: they are insufficiently focused on the future. In this paper, we theorize and demonstrate that encouraging input providers to think about the future leads them to produce more concrete developmental input. Across a large scale, preregistered field experiment (n = 27,432 comments) and two laboratory studies (n = 806), people provide more concrete and actionable developmental input when they are prompted to provide future-looking “advice” rather than “feedback,” a common method of soliciting input in organizations. The effect of soliciting advice on input concreteness was mediated by providers’ future focus. Moreover, in a follow-up study, such concrete input was assessed by independent raters as more useful. These findings highlight the role of temporal orientation in driving the content of developmental input. In doing so, our data suggest that individuals and organizations have the potential to promote higher-quality developmental input by attending to the temporal orientation that their input systems encourage. This paper was accepted by Isabel Fernandez-Mateo, organizations. Funding: Funding for the laboratory experiments was provided by Harvard Business School. Supplemental Material: The supplementary appendix and data files are available at https://doi.org/10.1287/mnsc.2022.03207 .