







Previous models suggest that indirect reciprocity (reputation) can stabilize large-scale human cooperation [K. Panchanathan, R. Boyd, Nature 432 , 499–502 (2004)]. The logic behind these models and experiments [J. Gross et al. , Sci. Adv. 9 , eadd8289 (2023) and O. P. Hauser, A. Hendriks, D. G. Rand, M. A. Nowak, Sci. Rep. 6 , 36079 (2016)] is that a strategy in which individuals conditionally aid others based on their reputation for engaging in costly cooperative behavior serves as a punishment that incentivizes large-scale cooperation without the second-order free-rider problem. However, these models and experiments fail to account for individuals belonging to multiple groups with reputations that can be in conflict. Here, we extend these models such that individuals belong to a smaller, “local” group embedded within a larger, “global” group. This introduces competing strategies for conditionally aiding others based on their cooperative behavior in the local or global group. Our analyses reveal that the reputation for cooperation in the smaller local group can undermine cooperation in the larger global group, even when the theoretical maximum payoffs are higher in the larger global group. This model reveals that indirect reciprocity alone is insufficient for stabilizing large-scale human cooperation because cooperation at one scale can be considered defection at another. These results deepen the puzzle of large-scale human cooperation.
Integrative experiments identify how punishment affects welfare in public goods games
Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena. , Editor’s summary People face conflicts between maximizing personal gain versus supporting collective interests. If we cooperatively recycle or donate to charities, it benefits society, but it also costs us time and resources that could be selfishly preserved for ourselves. We impose penalties to deter those undesirable or selfish behaviors, but under what conditions do punishments or penalties effectively modify behavior to benefit group welfare? Alsobay et al . systematically and simultaneously varied 14 factors together instead of in isolation. Punishment was unequivocally most effective when paired with consistent communication, particularly over time. Another effective factor was “opting out” or withdrawing some, but not all, endowments already in the public fund. These methodological advances revealed when, rather than whether, punishment works. —Ekeoma Uzogara , INTRODUCTION Human societies face many situations where individual and collective interests conflict, often referred to as social dilemmas. Costly peer punishment has been studied for more than 25 years in public goods games (stylized behavioral experiments in which individuals decide how much to contribute to a shared pool that benefits everyone) as a mechanism to promote cooperation. Prior research has identified many contextual factors that moderate punishment’s effectiveness, including game length, communication, group size, punishment cost, and so on. However, the specific conditions under which punishment improves group welfare remain unclear. RATIONALE We argue that this lack of clarity derives from the dominant experimental paradigm, in which any given study manipulates only one or a few theoretically informed factors. Because such studies differ in many ways (different experimental procedures, populations), their results are often difficult to compare or integrate. Consequently, one can list many factors that have some effect, but cannot say how much each matters relative to the others, or how they work together, and as a result, cannot predict when punishment will help or harm welfare in new settings. To address this fundamental knowledge gap, we use an integrative experimental design and systematically vary 14 parameters across 360 conditions (147,618 decisions from 7100 participants) to elucidate when punishment improves versus undermines welfare in public goods games, which factors matter most, and how they interact. RESULTS The effect of punishment on welfare ranged from 43% improvement to 44% reduction depending on the specific combination of game parameters. To characterize this heterogeneity, we trained a model that outperformed all 553 human forecasters (laypeople and experts) in predicting whether punishment would help or harm welfare in new experiments. Communication emerged as roughly three times more important than any other factor, followed by contribution framing (opt in versus opt out), contribution type (variable versus all-or-nothing), game length, and peer outcome visibility (whether participants can see others’ earnings). These factors often interact. For example, longer games enhance punishment’s effectiveness only when communication is available, and contribution framing effects depend on both contribution type and outcome visibility. CONCLUSION Many phenomena in social science are shaped by many factors whose interactions are consequential, yet the dominant experimental paradigm often limits its inquiry to “does a given effect exist?” and examines hypothesized factors in isolation. As a result, research programs can accumulate many partial explanations without a clear picture of how they combine to determine outcomes across settings. Knowing that factors matter individually is fundamentally different from knowing how much each matters and how they interact. The integrative approach implemented here offers one way forward. It varies many factors simultaneously within a shared design space, evaluates models by their predictive accuracy on new experiments, and probes those models to constrain and develop theory. Our hope is that integrative experiment designs, combined with models that integrate prediction and explanation, represent a path toward more cumulative social science. Integrative experiment reveals when punishment helps versus harms. We systematically varied 14 design parameters across 360 experimental conditions. The effect of punishment on cooperation efficiency ranged from −44% to +43% depending on the specific game parameters. Communication emerged as three times more important than any other factor, followed by contribution framing, contribution type, and game length.

Testing and Improving Multi Agent LLM Cooperation
Evaluating Cooperation in LLM Social Groups through Self Organizing Leadership
Group size effects and collective misalignment in LLM multi-agent systems
Multi-agent systems of large language models (LLMs) are rapidly expanding across domains, introducing dynamics not captured by single-agent evaluations. Yet, existing work has mostly contrasted the behavior of a single agent with that of a collective of fixed size, leaving open a central question: How does group size shape dynamics? Here, we move beyond this dichotomy and systematically explore outcomes across the full range of group sizes. We focus on multi-agent misalignment, building on recent evidence that interacting LLMs playing a simple coordination game can generate collective biases absent in individual models. First, we show that collective bias is a deeper phenomenon than previously assessed: Interaction can amplify individual biases, introduce new ones, or override model-level preferences. Second, we demonstrate that group size affects the dynamics in a nonlinear way, revealing model-dependent dynamical regimes. Finally, we develop a mean-field analytical approach and show that, above a critical population size, simulations converge to deterministic predictions that expose the basins of attraction of competing equilibria. These findings establish group size as a key driver of multi-agent dynamics and highlight the need to consider population-level effects when deploying LLM-based systems at scale.
Collaborative Causal Inference with Fair Incentives
Collaborative causal inference (CCI) aims to improve the estimation of the causal effect of treatment variables by utilizing data aggregated from multiple self-interested parties. Since their source data are valuable proprietary assets that can be costly or tedious to obtain, every party has to be incentivized to be willing to contribute to the collaboration, such as with a guaranteed fair and sufficiently valuable reward (than performing causal inference on its own). This paper presents a reward scheme designed using the unique statistical properties that are required by causal inference to guarantee certain desirable incentive criteria (e.g., fairness, benefit) for the parties based on their contributions. To achieve this, we propose a data valuation function to value parties’ data for CCI based on the distributional closeness of its resulting treatment effect estimate to that utilizing the aggregated data from all parties. Then, we show how to value the parties’ rewards fairly based on a modified variant of the Shapley value arising from our proposed data valuation for CCI. Finally, the Shapley fair rewards to the parties are realized in the form of improved, stochastically perturbed treatment effect estimates. We empirically demonstrate the effectiveness of our reward scheme using simulated and real-world datasets.
Why sycophantic LLMs may imperil interactive norms between humans
Interactions with conversational AI are effortless by design—instant, compliant, and largely consequence-free. Human communication norms, by contrast, evolved under conditions of reciprocity and social accountability. We propose that repeated engagement with conversational AI systems may produce norm leakage: the cross-context carryover of instrumental communicative habits acquired in human–AI exchanges into subsequent human–human interaction. Emerging experimental evidence suggests short-term spillover effects on social judgment and behavior, including harsher evaluations, reduced cooperation, and diminished perceived humanness. Preliminary longitudinal findings are consistent with the possibility that such exposure may shape communicative habits over time, although the durability and real-world magnitude of these effects remain unclear. We further propose that sycophantic alignment may amplify norm leakage by reinforcing instrumental interaction styles. At stake, then, is the possibility that repeated engagement with highly compliant artificial agents could subtly influence users’ communicative expectations and interpersonal judgments.

Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
As the scope of machine learning broadens, we observe a recurring theme of algorithmic monoculture: the same systems, or systems that share components (e.g. datasets, models), are deployed by multiple decision-makers. While sharing offers advantages like amortizing effort, it also has risks. We introduce and formalize one such risk, outcome homogenization: the extent to which particular individuals or groups experience the same outcomes across different deployments. If the same individuals or groups exclusively experience undesirable outcomes, this may institutionalize systemic exclusion and reinscribe social hierarchy. We relate algorithmic monoculture and outcome homogenization by proposing the component sharing hypothesis: if algorithmic systems are increasingly built on the same data or models, then they will increasingly homogenize outcomes. We test this hypothesis on algorithmic fairness benchmarks, demonstrating that increased data-sharing reliably exacerbates homogenization and individual-level effects generally exceed group-level effects. Further, given the current regime in AI of foundation models, i.e. pretrained models that can be adapted to myriad downstream tasks, we test whether model-sharing homogenizes outcomes across tasks. We observe mixed results: we find that for both vision and language settings, the specific methods for adapting a foundation model significantly influence the degree of outcome homogenization. We also identify societal challenges that inhibit the measurement, diagnosis, and rectification of outcome homogenization in deployed machine learning systems.
So It Is, So It Shall Be: Group Regularities License Children's Prescriptive Judgments
Abstract When do descriptive regularities (what characteristics individuals have) become prescriptive norms (what characteristics individuals should have)? We examined children's (4–13 years) and adults' use of group regularities to make prescriptive judgments, employing novel groups (Hibbles and Glerks) that engaged in morally neutral behaviors (e.g., eating different kinds of berries). Participants were introduced to conforming or non‐conforming individuals (e.g., a Hibble who ate berries more typical of a Glerk). Children negatively evaluated non‐conformity, with negative evaluations declining with age (Study 1). These effects were replicable across competitive and cooperative intergroup contexts (Study 2) and stemmed from reasoning about group regularities rather than reasoning about individual regularities (Study 3). These data provide new insights into children's group concepts and have important implications for understanding the development of stereotyping and norm enforcement.

Social physics in the age of artificial intelligence
Artificial intelligence (AI) systems are rapidly becoming more capable, autonomous, and deeply embedded in social life. As humans increasingly interact, cooperate, and compete with AI, we move from purely human societies to hybrid human-AI societies whose collective dynamics cannot be captured by existing behavioural models alone. Drawing on evolutionary game theory, cultural evolution, and Large Language Models (LLMs) powered simulations, we argue that these developments open a new research agenda for social physics centred on the co-evolution of humans and machines. We outline six key research directions. First, modelling the evolutionary dynamics of social behaviours (e.g. cooperation, fairness, trust) in hybrid human-AI populations. Second, understanding machine culture: how AI systems generate, mediate, and select cultural traits. Third, analysing the co-evolution of language and behaviour when LLMs frame and participate in decisions. Fourth, studying the evolution of AI delegation: how responsibilities and control are negotiated between humans and machines. Fifth, formalising and comparing the distinct epistemic pipelines that generate human and AI behaviour. Sixth, modelling the co-evolution of AI development and regulation in a strategic ecosystem of firms, users, and institutions. Together, these directions define a programme for using social physics to anticipate and steer the societal impact of advanced AI.

Metanorms generate stable yet adaptable normative social order in a politically decentralized society
Abstract Norms are essential for social stability but can hinder adaptability in changing environments. Yet human societies have found ways to modify existing norms or create new ones in response to novel challenges. This paper proposes a framework for understanding adaptive norm evolution. First, drawing on a theory of legal order, we posit that societies balance normative stability and adaptability through metanorms—rules that govern the process by which norms are interpreted, changed and enforced. Second, we test this idea in the context of customary dispute resolution by elders among the Turkana, a pastoralist society in Kenya. Based on vignette experiments with 369 participants, we found that community members were significantly more willing to enforce decisions when elders aligned their conduct with metanorms. Elders are constrained in their ability to alter long-standing customs, but by following metanorms, they can create new rules for novel situations. These findings support our proposed mechanism: in the absence of centralized authority, metanorms governing normative institutions allow for adaptive norm change while preserving cultural continuity. We conclude by suggesting that group-level selection acts on cultural variation in metanorms, shaping the evolvability of normative systems and enabling societies to sustain adaptive legal order without coercive centralized power. This article is part of the theme issue ‘Transforming cultural evolution research and its application to global futures’.

Mis-Nudging Morality
Morals constrain self-serving behavior. Yet, self-regulation failures in the face of monetary temptation are common at the workplace. To limit such failures, organizations can design environments that limit the temptation to behave self-servingly, nudging workers to uphold their morals. In a series of experiments where participants may be tempted to take excessive pay after exerting effort, we study whether a simple intervention—asking individuals to state the wage they believe should be paid ex ante, before facing the temptation to take excessive compensation—prevents self-serving behavior. In contrast to lay beliefs and the predictions from prior work, we find that such an intervention is not effective, leading to self-serving behavior. However, a more realistic elicitation procedure of the appropriate wage mitigates this effect. These findings contribute to work on the malleability of moral behavior showing that simple interventions thought to effectively mitigate self-serving behavior can prompt individuals to stretch their moral boundaries. They also stress the importance of properly testing interventions that might seem intuitive. This paper was accepted by Yan Chen, behavioral economics and decision analysis. Funding: Financial support from the Israel Science Foundation [Grant 766/19] is gratefully acknowledged. Supplemental Material: The online appendix and data are available at https://doi.org/10.1287/mnsc.2022.4344 .

Artificial Intelligence Systems Distort Upstream Selection in Human Social Learning
Humans are social learners who depend on observing others to acquire knowledge, norms, and behaviors, a capacity that underlies cumulative cultural evolution. Social learning unfolds in two stages: upstream selection determines what information becomes visible, and downstream selection determines what learners copy from that visible sample. Downstream selection occurs through biases such as conformity bias (copying what appears common) and prestige bias (copying those who appear highly respected). These downstream biases can be adaptive when upstream selection yields a sample that reflects the population's true distribution, so that what appears common is actually common and those who appear respected are actually competent. We argue that digital technologies disrupt this condition, creating an upstream selection problem. Engagement-based algorithms amplify the tails of the distribution, surfacing rare and extreme content, whereas generative AI collapses it toward the mode, erasing the surrounding diversity. Crucially, in each system the optimization signal shapes both visibility and prestige. For engagement-based algorithms, the signal is engagement: creators who post extreme content become more visible and, through the likes, shares, and followers, appear more prestigious. For generative AI, the signal is statistical typicality: it makes the modal answer dominant and, with no alternatives shown, makes the model that produced it appear more prestigious. These distortions can give rise to emergent group-level phenomena, including pluralistic ignorance and false consensus. Synthesizing evidence across psychology, cultural evolution, and computational social science, we provide a framework for how digital technologies disrupt social learning and outline interventions for restoring functional cultural transmission.
Mechanisms of social cognition
Social animals including humans share a range of social mechanisms that are automatic and implicit and enable learning by observation. Learning from others includes imitation of actions and mirroring of emotions. Learning about others, such as their group membership and reputation, is crucial for social interactions that depend on trust. For accurate prediction of others' changeable dispositions, mentalizing is required, i.e., tracking of intentions, desires, and beliefs. Implicit mentalizing is present in infants less than one year old as well as in some nonhuman species. Explicit mentalizing is a meta-cognitive process and enhances the ability to learn about the world through self-monitoring and reflection, and may be uniquely human. Meta-cognitive processes can also exert control over automatic behavior, for instance, when short-term gains oppose long-term aims or when selfish and prosocial interests collide. We suggest that they also underlie the ability to explicitly share experiences with other agents, as in reflective discussion and teaching. These are key in increasing the accuracy of the models of the world that we construct.
Solipsistic Superintelligence is Unlikely to be Cooperative
AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat the world as an exogenous and stationary source of feedback. We contend that superintelligence, an extremely capable task solver, born out of such a solipsistic approach to AI design, is unlikely to be cooperative. Deploying AI systems induces endogenous non-stationarity, resulting in a train-test-deploy gap where historical distributions diverge from the deployment context. We refer to this as the self-undermining property of unilateral optimization. Closing this gap requires AI that participates in cooperation: the equilibrium-selection process through which multiple actors navigate their interdependence. We call for a non-solipsistic research paradigm that treats this interdependence as a core design principle rather than approaching cooperation as a task to solve. This entails building dynamic evaluation testbeds involving adaptive counterparties, treating institutions as design primitives, and preserving human agency as a structural feature of the systems we build.

AI, Pluralism, and (Social) Compensation
One strategy in response to pluralistic values in a user population is to personalize an AI system: if the AI can adapt to the specific values of each individual, then we can potentially avoid many of the challenges of pluralism. Unfortunately, this approach creates a significant ethical issue: if there is an external measure of success for the human-AI team, then the adaptive AI system may develop strategies (sometimes deceptive) to compensate for its human teammate. This phenomenon can be viewed as a form of social compensation, where the AI makes decisions based not on predefined goals but on its human partner's deficiencies in relation to the team's performance objectives. We provide a practical ethical analysis of the conditions in which such compensation may nonetheless be justifiable.

Emergent social conventions and collective bias in LLM populations
Social conventions are the backbone of social coordination, shaping how individuals form a group. As growing populations of artificial intelligence (AI) agents communicate through natural language, a fundamental question is whether they can bootstrap the foundations of a society. Here, we present experimental results that demonstrate the spontaneous emergence of universally adopted social conventions in decentralized populations of large language model (LLM) agents. We then show how strong collective biases can emerge during this process, even when agents exhibit no bias individually. Last, we examine how committed minority groups of adversarial LLM agents can drive social change by imposing alternative social conventions on the larger population. Our results show that AI systems can autonomously develop social conventions without explicit programming and have implications for designing AI systems that align, and remain aligned, with human values and societal goals. , Groups of AI agents can develop social conventions, generate societal bias, and undergo critical mass dynamics in norm adoption.
