







Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena. , Editor’s summary People face conflicts between maximizing personal gain versus supporting collective interests. If we cooperatively recycle or donate to charities, it benefits society, but it also costs us time and resources that could be selfishly preserved for ourselves. We impose penalties to deter those undesirable or selfish behaviors, but under what conditions do punishments or penalties effectively modify behavior to benefit group welfare? Alsobay et al . systematically and simultaneously varied 14 factors together instead of in isolation. Punishment was unequivocally most effective when paired with consistent communication, particularly over time. Another effective factor was “opting out” or withdrawing some, but not all, endowments already in the public fund. These methodological advances revealed when, rather than whether, punishment works. —Ekeoma Uzogara , INTRODUCTION Human societies face many situations where individual and collective interests conflict, often referred to as social dilemmas. Costly peer punishment has been studied for more than 25 years in public goods games (stylized behavioral experiments in which individuals decide how much to contribute to a shared pool that benefits everyone) as a mechanism to promote cooperation. Prior research has identified many contextual factors that moderate punishment’s effectiveness, including game length, communication, group size, punishment cost, and so on. However, the specific conditions under which punishment improves group welfare remain unclear. RATIONALE We argue that this lack of clarity derives from the dominant experimental paradigm, in which any given study manipulates only one or a few theoretically informed factors. Because such studies differ in many ways (different experimental procedures, populations), their results are often difficult to compare or integrate. Consequently, one can list many factors that have some effect, but cannot say how much each matters relative to the others, or how they work together, and as a result, cannot predict when punishment will help or harm welfare in new settings. To address this fundamental knowledge gap, we use an integrative experimental design and systematically vary 14 parameters across 360 conditions (147,618 decisions from 7100 participants) to elucidate when punishment improves versus undermines welfare in public goods games, which factors matter most, and how they interact. RESULTS The effect of punishment on welfare ranged from 43% improvement to 44% reduction depending on the specific combination of game parameters. To characterize this heterogeneity, we trained a model that outperformed all 553 human forecasters (laypeople and experts) in predicting whether punishment would help or harm welfare in new experiments. Communication emerged as roughly three times more important than any other factor, followed by contribution framing (opt in versus opt out), contribution type (variable versus all-or-nothing), game length, and peer outcome visibility (whether participants can see others’ earnings). These factors often interact. For example, longer games enhance punishment’s effectiveness only when communication is available, and contribution framing effects depend on both contribution type and outcome visibility. CONCLUSION Many phenomena in social science are shaped by many factors whose interactions are consequential, yet the dominant experimental paradigm often limits its inquiry to “does a given effect exist?” and examines hypothesized factors in isolation. As a result, research programs can accumulate many partial explanations without a clear picture of how they combine to determine outcomes across settings. Knowing that factors matter individually is fundamentally different from knowing how much each matters and how they interact. The integrative approach implemented here offers one way forward. It varies many factors simultaneously within a shared design space, evaluates models by their predictive accuracy on new experiments, and probes those models to constrain and develop theory. Our hope is that integrative experiment designs, combined with models that integrate prediction and explanation, represent a path toward more cumulative social science. Integrative experiment reveals when punishment helps versus harms. We systematically varied 14 design parameters across 360 experimental conditions. The effect of punishment on cooperation efficiency ranged from −44% to +43% depending on the specific game parameters. Communication emerged as three times more important than any other factor, followed by contribution framing, contribution type, and game length.
Indirect reciprocity undermines indirect reciprocity destabilizing large-scale cooperation
Previous models suggest that indirect reciprocity (reputation) can stabilize large-scale human cooperation [K. Panchanathan, R. Boyd, Nature 432 , 499–502 (2004)]. The logic behind these models and experiments [J. Gross et al. , Sci. Adv. 9 , eadd8289 (2023) and O. P. Hauser, A. Hendriks, D. G. Rand, M. A. Nowak, Sci. Rep. 6 , 36079 (2016)] is that a strategy in which individuals conditionally aid others based on their reputation for engaging in costly cooperative behavior serves as a punishment that incentivizes large-scale cooperation without the second-order free-rider problem. However, these models and experiments fail to account for individuals belonging to multiple groups with reputations that can be in conflict. Here, we extend these models such that individuals belong to a smaller, “local” group embedded within a larger, “global” group. This introduces competing strategies for conditionally aiding others based on their cooperative behavior in the local or global group. Our analyses reveal that the reputation for cooperation in the smaller local group can undermine cooperation in the larger global group, even when the theoretical maximum payoffs are higher in the larger global group. This model reveals that indirect reciprocity alone is insufficient for stabilizing large-scale human cooperation because cooperation at one scale can be considered defection at another. These results deepen the puzzle of large-scale human cooperation.

The Effects of Financial Incentives in Experiments: A Review and Capital-Labor-Production Framework
We review 74 experiments with no, low, or high performance-based financial incentives. The modal result has no effect on mean performance (though variance is usually reduced by higher payment). Higher incentive does improve performance often, typically judgment tasks that are responsive to better effort. Incentives also reduce “presentation” effects (e.g., generosity and risk-seeking). Incentive effects are comparable to effects of other variables, particularly “cognitive capital” and task “production” demands, and interact with those variables, so a narrow-minded focus on incentives alone is misguided. We also note that no replicated study has made rationality violations disappear purely by raising incentives.
How Field Experiments in Economics Can Complement Psychological Research on Judgment Biases
This review summarizes results of field experiments examining individual behaviors across several market settings—from open-air markets to rideshare markets to tax-compliance markets—where people sort themselves into market roles wherein they make consequential decisions. Using three distinct examples from my own research on the endowment effect, left-digit bias, and omission bias, I showcase how field experiments can help researchers understand mediators, heterogeneity, and causal moderation involved in judgment biases in the field. In this manner, the review highlights that economic field experiments can serve an invaluable intellectual role alongside traditional laboratory research.

Mis-Nudging Morality
Morals constrain self-serving behavior. Yet, self-regulation failures in the face of monetary temptation are common at the workplace. To limit such failures, organizations can design environments that limit the temptation to behave self-servingly, nudging workers to uphold their morals. In a series of experiments where participants may be tempted to take excessive pay after exerting effort, we study whether a simple intervention—asking individuals to state the wage they believe should be paid ex ante, before facing the temptation to take excessive compensation—prevents self-serving behavior. In contrast to lay beliefs and the predictions from prior work, we find that such an intervention is not effective, leading to self-serving behavior. However, a more realistic elicitation procedure of the appropriate wage mitigates this effect. These findings contribute to work on the malleability of moral behavior showing that simple interventions thought to effectively mitigate self-serving behavior can prompt individuals to stretch their moral boundaries. They also stress the importance of properly testing interventions that might seem intuitive. This paper was accepted by Yan Chen, behavioral economics and decision analysis. Funding: Financial support from the Israel Science Foundation [Grant 766/19] is gratefully acknowledged. Supplemental Material: The online appendix and data are available at https://doi.org/10.1287/mnsc.2022.4344 .

Putting nudges in perspective
Conventional economic policy focuses on ‘economic’ solutions (e.g. taxes, incentives, regulation) to problems caused by market-level factors such as externalities, misaligned incentives and information asymmetries. By contrast, ‘nudges’ provide behavioural solutions to problems that have generally been assumed to originate from limitations in human decision making, such as present bias. While policy-makers have good reason for exploiting the power of nudges, we argue that these extremes leave open a large space of policy options that have received less attention in the academic literature. First, there is no reason that solution and problem need have the same theoretical basis: there are promising behavioural solutions to problems that have causes that are well explained by traditional economics, and conventional economic solutions often offer the best line of attack on problems of behavioural origin. Second, there is a wide range of hybrid policy actions with both economic and behavioural components (e.g. framing a tax or incentive in a specific way), and there exist many societal problems – perhaps the majority – that arise from both economic and behavioural factors (e.g. firms’ exploitation of consumers’ behavioural biases). This paper aims to remind policy-makers that behavioural economics can influence policy in a variety of ways, of which nudges are the most prominent but not necessarily the most powerful.

The Challenge of Understanding What Users Want: Inconsistent Preferences and Engagement Optimization
Online platforms have a wealth of data, run countless experiments, and use industrial-scale algorithms to optimize user experience. Despite this, many users seem to regret the time they spend on these platforms. One possible explanation is that incentives are misaligned: platforms are not optimizing for user happiness. We suggest the problem runs deeper, transcending the specific incentives of any particular platform, and instead stems from a mistaken foundational assumption. To understand what users want, platforms look at what users do. This is a kind of revealed-preference assumption that is ubiquitous in the way user models are built. Yet research has demonstrated, and personal experience affirms, that we often make choices in the moment that are inconsistent with what we actually want. The behavioral economics and psychology literatures suggest, for example, that we can choose mindlessly or that we can be too myopic in our choices, behaviors that feel entirely familiar on online platforms. In this work, we develop a model of media consumption where users have inconsistent preferences. We consider a platform which wants to maximize user utility, but only observes behavioral data in the form of the user’s engagement. We show how our model of users’ preference inconsistencies produces phenomena that are familiar from everyday experience but difficult to capture in traditional user interaction models. These phenomena include users who have long sessions on a platform but derive very little utility from it, and platform changes that steadily raise user engagement before abruptly causing users to go “cold turkey” and quit. A key ingredient in our model is a formulation for how platforms determine what to show users: they optimize over a large set of potential content (the content manifold) parametrized by underlying features of the content. Whether improving engagement improves user welfare depends on the direction of movement in the content manifold: For certain directions of change, increasing engagement makes users less happy, whereas in other directions on the same manifold, increasing engagement makes users happier. We provide a characterization of the structure of content manifolds for which increasing engagement fails to increase user utility. By linking these effects to abstractions of platform design choices, our model thus creates a theoretical framework and vocabulary in which to explore interactions between design, behavioral science, and social media. This paper was accepted by Yan Chen, behavioral economics and decision analysis. Funding: This work was supported by the Vannevar Bush Faculty Fellowship and Multidisciplinary University Research Initiative [Grant W911NF-19-0217]. Supplemental Material: The online appendices are available at https://doi.org/10.1287/mnsc.2022.03683 .

Mechanism Experiments and Policy Evaluations
Randomized controlled trials are increasingly used to evaluate policies. How can we make these experiments as useful as possible for policy purposes? We argue greater use should be made of experiments that identify the behavioral mechanisms that are central to clearly specified policy questions, what we call "mechanism experiments." These types of experiments can be of great policy value even if the intervention that is tested (or its setting) does not correspond exactly to any realistic policy option.
Underspecified Human Decision Experiments Considered Harmful
Decision-making with information displays is a key focus of research in areas like human-AI collaboration and data visualization. However, what constitutes a decision problem, and what is required for an experiment to conclude that decisions are flawed, remain imprecise. We present a widely applicable definition of a decision problem synthesized from statistical decision theory and information economics. We claim that to attribute loss in human performance to bias, an experiment must provide the information that a rational agent would need to identify the normative decision. We evaluate whether recent empirical research on AI-assisted decisions achieves this standard. We find that only 10 (26%) of 39 studies that claim to identify biased behavior presented participants with sufficient information to make this claim in at least one treatment condition. We motivate the value of studying well-defined decision problems by describing a characterization of performance losses they allow to be conceived.

Monetary incentives, what are they good for?
This paper is a critical reflection on the use of monetary incentives in economic experiments. The argument is that incentives have their effect through their influence on one or more of three fact...

Characterizing the causes, dynamics, and consequences of choice deferral
The fact that people often avoid making decisions is well known, and past research has helped to identify some of the conditions and reasons for doing so. For instance, people may forgo choices between bad options because they prefer not to end up with one of those options. It is much less clear when and why people avoid choosing in cases where they eventually will have to make a given decision. To study such instances of choice deferral, we presented participants with a series of choices, and for each choice they were allowed to either choose immediately or defer the decision until later in the experiment. Across six experiments and three choice domains (choices among consumer goods, artwork, and political candidates), we find that the strongest predictor of choice deferral is the overall value of a given set of options, with relative value (i.e., how hard it is to identify the best option) counterintuitively playing a smaller role. We show that the influence of overall value on choice deferral can be accounted for by a dynamic decision model according to which participants appraise the option set relative to a criterion before deciding whether to choose or defer, comparing this to a previous model whereby participants make such a decision based on a predetermined decision time limit. We further reveal that the influence of overall value on choice deferral is determined by how congruent options are with a given choice goal (choose-best or choose-worst) rather than simply how bad those options are. Collectively, our findings shed new light on how people decide to put off the inevitable.
Cheating in the Lab Predicts Fraud in the Field: An Experiment in Public Transportation
We conduct an artefactual field experiment using a diversified sample of passengers of public transportation to study attitudes toward dishonesty. We find that the diversity of behavior in terms of (dis)honesty in laboratory tasks and in the field correlate. Moreover, individuals who have just been fined in the field behave more honestly in the lab than the other fare dodgers, except when context is introduced. Overall, we show that simple tests of dishonesty in the lab can predict moral firmness in life, although fraudsters who care about social image cheat less when behavior can be verified ex post by the experimenter. Data and the online appendix are available at https://doi.org/10.1287/mnsc.2016.2616 . This paper was accepted by Uri Gneezy, behavioral economics.

Anticipation and Choice Heuristics in the Dynamic Consumption of Pain Relief
Humans frequently need to allocate resources across multiple time-steps. Economic theory proposes that subjects do so according to a stable set of intertemporal preferences, but the computational demands of such decisions encourage the use of formally less competent heuristics. Few empirical studies have examined dynamic resource allocation decisions systematically. Here we conducted an experiment involving the dynamic consumption over approximately 15 minutes of a limited budget of relief from moderately painful stimuli. We had previously elicited the participants’ time preferences for the same painful stimuli in one-off choices, allowing us to assess self-consistency. Participants exhibited three characteristic behaviors: saving relief until the end, spreading relief across time, and early spending, of which the last was markedly less prominent. The likelihood that behavior was heuristic rather than normative is suggested by the weak correspondence between one-off and dynamic choices. We show that the consumption choices are consistent with a combination of simple heuristics involving early-spending, spreading or saving of relief until the end, with subjects predominantly exhibiting the last two.
Small Probabilistic Discounts Stimulate Spending: Pain of Paying in Price Promotions
AbstractWe find that small probabilistic price promotions effectively stimulate demand, even more so than comparable fixed price promotions (e.g., “1% chance it’s free” vs. “1% off,” respectively), because they more effectively reduce the pain of paying. In three field experiments at a grocer, we exogenously and endogenously manipulated the salience of pain of paying via elicitation timing (e.g., at entrance or checkout) and payment method (i.e., cash/debit cards or credit cards). This modulated the attractiveness of probabilistic discounts and their ability to stimulate spending. Shoppers paying with cash or debit cards, for example, spent 54% more if they received a 1% probabilistic discount than a 1% fixed discount (experiment 2). A fourth experiment showed that consumers’ sensitivity to pain of paying modulates the greater comparative efficacy of small probabilistic than fixed discounts. More broadly, the results elucidate a novel affective route through which price promotions stimulate demand––pain of paying.

You’ve Got Mail: A Randomized Field Experiment on Tax Evasion
We report from a large-scale randomized field experiment conducted on a unique sample of more than 15,000 taxpayers in Norway who were likely to have misreported their foreign income. By randomly manipulating a letter from the tax authorities, we cleanly identify that moral suasion and the perceived detection probability play a crucial role in shaping taxpayer behavior. The moral letter mainly works on the intensive margin, while the detection letter has a strong effect on the extensive margin. We further show that only the detection letter has long-term effects on tax compliance. This paper was accepted by Yan Chen, behavioral economics.

Unwillingness to pay for privacy: A field experiment
We measure willingness to pay for privacy in a field experiment. Participants bought at most one DVD from one of two competing online stores. One store consistently required more sensitive personal data than the other, but otherwise the stores were identical. In one treatment, DVDs were one Euro cheaper at the store requesting more personal information, and almost all buyers chose the cheaper store. Surprisingly, in the second treatment when prices were identical, participants bought from both shops equally often.
Do people like financial nudges?
Do people like financial nudges? To answer that question we conducted a pre-registered survey presenting people with 36 hypothetical scenarios describing financial interventions. We varied levels of transparency (i.e., explaining how the interventions worked), framing (interventions framed in terms of spending, or saving), and ‘System’ (interventions could target either System 1 or System 2). Participants were a random sample of 2,100 people drawn from a representative Australian population. All financial interventions were tested across six dependent variables: approval, benefit, ethics, manipulation, the likelihood of use, as well as the likelihood of use if the intervention were to be proposed by a bank. Results indicate that people generally approve of financial interventions, rating them as neutral to positive across all dependent variables (except for manipulation, which was reverse coded). We find effects of framing and System. People have strong and significant preferences for System 2 interventions, and interventions framed in terms of savings. Transparency was not found to have a significant impact on how people rate financial interventions. Financial interventions continue to be rated positive, regardless of the messenger. Looking at demographics, we find that participants who were female, younger, living in metro areas and earning higher incomes were most likely to favor financial interventions, and this effect is especially strong for those aged under 45. We discuss the implications for these results as applied to the financial sector.
