







Incentivized choice experiments are a key approach to measuring preferences in economics but are also costly. Survey measures are a low-cost alternative but can suffer from additional forms of meas...
Characterizing the causes, dynamics, and consequences of choice deferral
The fact that people often avoid making decisions is well known, and past research has helped to identify some of the conditions and reasons for doing so. For instance, people may forgo choices between bad options because they prefer not to end up with one of those options. It is much less clear when and why people avoid choosing in cases where they eventually will have to make a given decision. To study such instances of choice deferral, we presented participants with a series of choices, and for each choice they were allowed to either choose immediately or defer the decision until later in the experiment. Across six experiments and three choice domains (choices among consumer goods, artwork, and political candidates), we find that the strongest predictor of choice deferral is the overall value of a given set of options, with relative value (i.e., how hard it is to identify the best option) counterintuitively playing a smaller role. We show that the influence of overall value on choice deferral can be accounted for by a dynamic decision model according to which participants appraise the option set relative to a criterion before deciding whether to choose or defer, comparing this to a previous model whereby participants make such a decision based on a predetermined decision time limit. We further reveal that the influence of overall value on choice deferral is determined by how congruent options are with a given choice goal (choose-best or choose-worst) rather than simply how bad those options are. Collectively, our findings shed new light on how people decide to put off the inevitable.
A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
Environments built for people are increasingly operated by a new class of economic actors: LLM-powered software agents making decisions on our behalf. These decisions range from our purchases to travel plans to medical treatment selection. Current evaluations of these agents largely focus on task competence, but we argue for a deeper assessment: how these agents choose when faced with realistic decisions. We introduce ABxLab, a framework for systematically probing agentic choice through controlled manipulations of option attributes and persuasive cues. We apply this to a realistic web-based shopping environment, where we vary prices, ratings, and psychological nudges, all of which are factors long known to shape human choice. We find that agent decisions shift predictably and substantially in response, revealing that agents are strongly biased choosers even without being subject to the cognitive constraints that shape human biases. This susceptibility reveals both risk and opportunity: risk, because agentic consumers may inherit and amplify human biases; opportunity, because consumer choice provides a powerful testbed for a behavioral science of AI agents, just as it has for the study of human behavior. We release our framework as an open benchmark for rigorous, scalable evaluation of agent decision-making.

Adult age differences in monetary decisions with real and hypothetical reward
Abstract Age differences in monetary decisions may emerge because younger and older adults perceive the value of outcomes differently. Yet, age‐differential effects of monetary rewards on decisions are not well understood. Most laboratory studies on aging and decision making have used scenarios in which rewards were merely hypothetical (decisions did not have any real consequences) or in which only small amounts of money were at stake. In the current study, we compared younger adults' (20–29 years) and older adults' (61–82 years) decisions in probabilistic choice problems with real or hypothetical rewards. Decision‐contingent rewards were in a typical range of previous studies (gains of up to ~4.25 USD) or substantially scaled up (gains of up to ~85 USD per participant). Reward type (real vs. hypothetical) affected decision quality, including value maximization, switching between options, and dominance violations (choices of an option that was inferior to another option in all respects). Decision quality was markedly better with real than hypothetical rewards in older adults and correlated with numeracy in both age groups. However, we found no evidence that reward type affected people's risk preferences. Overall, the findings portray a fairly positive picture regarding the use of hypothetical scenarios to assess preferences: With carefully prepared instructions, people from different age groups indicate preferences in hypothetical scenarios that match their decisions with real and much higher rewards. One advantage of using real rewards is that they help to reduce decision noise.

Present-Biased Preferences and Credit Card Borrowing
Some individuals borrow extensively on their credit cards. This paper tests whether present-biased time preferences correlate with credit card borrowing. In a field study, we elicit individual time preferences with incentivized choice experiments, and match resulting time preference measures to individual credit reports and annual tax returns. The results indicate that present-biased individuals are more likely to have credit card debt, and to have significantly higher amounts of credit card debt, controlling for disposable income, other socio-demographics, and credit constraints. (JEL D12, D14, D91)
A Generalizable Scale of Propensity to Plan: The Long and the Short of Planning for Time and for Money
Abstract. Planning has pronounced effects on consumer behavior and intertemporal choice. We develop a six-item scale measuring individual differences in pr

Behavioral Impediments to Valuing Annuities: Complexity and Choice Bracketing
Abstract This paper examines two behavioral factors that diminish people's ability to value a lifetime income stream or annuity, drawing on a randomized experiment with about 4,000 adults in a U.S. nationally representative sample. We find that increasing the complexity of the annuity choice reduces respondents' ability to value the annuity, measured by the difference between the sell and buy values they assign to the annuity. When we limit narrow choice bracketing by inducing people to think first about how quickly or slowly to spend down assets in retirement, their ability to value an annuity increases.

Complexity and biases
We examine experimentally how complexity affects decision-making, when individuals choose among different products with varying benefits and costs. We find that complexity in costs leads to choosing a high-benefit product, with high costs and overall lower payoffs. In contrast, when complexity is in the benefits of the product, we cannot reject the hypothesis of random mistakes. We also examine the role of heterogeneous complexity. We find that individuals still (mistakenly) choose the high-benefit but costly product, even if cheaper and simple products are available. Our results suggest that salience is a main driver of choices under different forms of complexity.

Integrative experiments identify how punishment affects welfare in public goods games
Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena. , Editor’s summary People face conflicts between maximizing personal gain versus supporting collective interests. If we cooperatively recycle or donate to charities, it benefits society, but it also costs us time and resources that could be selfishly preserved for ourselves. We impose penalties to deter those undesirable or selfish behaviors, but under what conditions do punishments or penalties effectively modify behavior to benefit group welfare? Alsobay et al . systematically and simultaneously varied 14 factors together instead of in isolation. Punishment was unequivocally most effective when paired with consistent communication, particularly over time. Another effective factor was “opting out” or withdrawing some, but not all, endowments already in the public fund. These methodological advances revealed when, rather than whether, punishment works. —Ekeoma Uzogara , INTRODUCTION Human societies face many situations where individual and collective interests conflict, often referred to as social dilemmas. Costly peer punishment has been studied for more than 25 years in public goods games (stylized behavioral experiments in which individuals decide how much to contribute to a shared pool that benefits everyone) as a mechanism to promote cooperation. Prior research has identified many contextual factors that moderate punishment’s effectiveness, including game length, communication, group size, punishment cost, and so on. However, the specific conditions under which punishment improves group welfare remain unclear. RATIONALE We argue that this lack of clarity derives from the dominant experimental paradigm, in which any given study manipulates only one or a few theoretically informed factors. Because such studies differ in many ways (different experimental procedures, populations), their results are often difficult to compare or integrate. Consequently, one can list many factors that have some effect, but cannot say how much each matters relative to the others, or how they work together, and as a result, cannot predict when punishment will help or harm welfare in new settings. To address this fundamental knowledge gap, we use an integrative experimental design and systematically vary 14 parameters across 360 conditions (147,618 decisions from 7100 participants) to elucidate when punishment improves versus undermines welfare in public goods games, which factors matter most, and how they interact. RESULTS The effect of punishment on welfare ranged from 43% improvement to 44% reduction depending on the specific combination of game parameters. To characterize this heterogeneity, we trained a model that outperformed all 553 human forecasters (laypeople and experts) in predicting whether punishment would help or harm welfare in new experiments. Communication emerged as roughly three times more important than any other factor, followed by contribution framing (opt in versus opt out), contribution type (variable versus all-or-nothing), game length, and peer outcome visibility (whether participants can see others’ earnings). These factors often interact. For example, longer games enhance punishment’s effectiveness only when communication is available, and contribution framing effects depend on both contribution type and outcome visibility. CONCLUSION Many phenomena in social science are shaped by many factors whose interactions are consequential, yet the dominant experimental paradigm often limits its inquiry to “does a given effect exist?” and examines hypothesized factors in isolation. As a result, research programs can accumulate many partial explanations without a clear picture of how they combine to determine outcomes across settings. Knowing that factors matter individually is fundamentally different from knowing how much each matters and how they interact. The integrative approach implemented here offers one way forward. It varies many factors simultaneously within a shared design space, evaluates models by their predictive accuracy on new experiments, and probes those models to constrain and develop theory. Our hope is that integrative experiment designs, combined with models that integrate prediction and explanation, represent a path toward more cumulative social science. Integrative experiment reveals when punishment helps versus harms. We systematically varied 14 design parameters across 360 experimental conditions. The effect of punishment on cooperation efficiency ranged from −44% to +43% depending on the specific game parameters. Communication emerged as three times more important than any other factor, followed by contribution framing, contribution type, and game length.

How Field Experiments in Economics Can Complement Psychological Research on Judgment Biases
This review summarizes results of field experiments examining individual behaviors across several market settings—from open-air markets to rideshare markets to tax-compliance markets—where people sort themselves into market roles wherein they make consequential decisions. Using three distinct examples from my own research on the endowment effect, left-digit bias, and omission bias, I showcase how field experiments can help researchers understand mediators, heterogeneity, and causal moderation involved in judgment biases in the field. In this manner, the review highlights that economic field experiments can serve an invaluable intellectual role alongside traditional laboratory research.

Behavioural economics, consumer behaviour and consumer policy: state of the art
Counter to the traditional assumption of neoclassical economics that individuals are rational Homo oeconomici that always seek to maximize their utility and follow their ‘true’ preferences, research in behavioural economics has demonstrated that people's judgements and decisions are often subject to systematic biases and heuristics, and are strongly dependent on the context of the decision. In this article, we briefly review the transition of research from neoclassical economics to behavioural economics, and discuss how the latter has influenced research in consumer behaviour and consumer policy. In particular, we discuss the impacts of key principles such as status quo bias, the endowment effect, mental accounting and the sunk-cost effect, other heuristics and biases related to availability, salience, the anchoring effect and simplicity rules, as well as the effects of other supposedly irrelevant factors such as music, temperature and physical markers on consumers’ decisions. These principles not only add significantly to research on consumer behaviour – they also offer readily available practical implications for consumer policy to nudge behaviour in beneficial directions in consumption domains including financial decision making, product choice, healthy eating and sustainable consumption.

Small Probabilistic Discounts Stimulate Spending: Pain of Paying in Price Promotions
AbstractWe find that small probabilistic price promotions effectively stimulate demand, even more so than comparable fixed price promotions (e.g., “1% chance it’s free” vs. “1% off,” respectively), because they more effectively reduce the pain of paying. In three field experiments at a grocer, we exogenously and endogenously manipulated the salience of pain of paying via elicitation timing (e.g., at entrance or checkout) and payment method (i.e., cash/debit cards or credit cards). This modulated the attractiveness of probabilistic discounts and their ability to stimulate spending. Shoppers paying with cash or debit cards, for example, spent 54% more if they received a 1% probabilistic discount than a 1% fixed discount (experiment 2). A fourth experiment showed that consumers’ sensitivity to pain of paying modulates the greater comparative efficacy of small probabilistic than fixed discounts. More broadly, the results elucidate a novel affective route through which price promotions stimulate demand––pain of paying.

Choice Bracketing
When making many choices, a person can broadly bracket them by assessing the consequences of all of them taken together, or narrowly bracket them by making each choice in isolation. We integrate research conducted in a wide range of decision contexts which shows that choice bracketing is an important determinant of behavior. Because broad bracketing allows people to take into account all the consequences of their actions, it generally leads to choices that yield higher utility. The evidence that we review, however, shows that people often fail to bracket broadly when it would be feasible for them to do so. In addition to documenting the diverse effects of bracketing, we also discuss factors that determine whether people bracket narrowly or broadly. We conclude with a discussion of normative aspects of bracketing and argue that there are some situations in which narrower bracketing results in superior decision making.

Context-Dependent Drivers of Discretionary Debt Decisions: Explaining Willingness to Borrow for Experiential Purchases
Abstract Mental accounting research suggests that consumers prefer borrowing for longer-lasting purchases in order to receive benefits from the purchases as they pay for them. In contrast, two sets of archival data and five lab studies show that consumers are more willing to borrow for experiential versus material purchases, even though experiential purchases tend to have a shorter physical duration. Further, framing the same purchase as more experiential than material increases willingness to borrow. This effect occurs because purchase timing is more important for experiential purchases—a function of consumers’ aversion to missing out on planned consumption. Thus, we moderate the proposed effect by varying whether the borrowing decision impacts planned consumption. Other differences between material and experiential purchases, such as scarcity or expected happiness, cannot similarly explain our results. Moreover, our conceptualization allows us to reconcile the apparent contradiction between the previous and current research by examining the relative impact of purchase-timing importance and payment-benefit duration matching in different contexts (i.e., “purchasing” and “source-of-funding” decisions).

The Challenge of Understanding What Users Want: Inconsistent Preferences and Engagement Optimization
Online platforms have a wealth of data, run countless experiments, and use industrial-scale algorithms to optimize user experience. Despite this, many users seem to regret the time they spend on these platforms. One possible explanation is that incentives are misaligned: platforms are not optimizing for user happiness. We suggest the problem runs deeper, transcending the specific incentives of any particular platform, and instead stems from a mistaken foundational assumption. To understand what users want, platforms look at what users do. This is a kind of revealed-preference assumption that is ubiquitous in the way user models are built. Yet research has demonstrated, and personal experience affirms, that we often make choices in the moment that are inconsistent with what we actually want. The behavioral economics and psychology literatures suggest, for example, that we can choose mindlessly or that we can be too myopic in our choices, behaviors that feel entirely familiar on online platforms. In this work, we develop a model of media consumption where users have inconsistent preferences. We consider a platform which wants to maximize user utility, but only observes behavioral data in the form of the user’s engagement. We show how our model of users’ preference inconsistencies produces phenomena that are familiar from everyday experience but difficult to capture in traditional user interaction models. These phenomena include users who have long sessions on a platform but derive very little utility from it, and platform changes that steadily raise user engagement before abruptly causing users to go “cold turkey” and quit. A key ingredient in our model is a formulation for how platforms determine what to show users: they optimize over a large set of potential content (the content manifold) parametrized by underlying features of the content. Whether improving engagement improves user welfare depends on the direction of movement in the content manifold: For certain directions of change, increasing engagement makes users less happy, whereas in other directions on the same manifold, increasing engagement makes users happier. We provide a characterization of the structure of content manifolds for which increasing engagement fails to increase user utility. By linking these effects to abstractions of platform design choices, our model thus creates a theoretical framework and vocabulary in which to explore interactions between design, behavioral science, and social media. This paper was accepted by Yan Chen, behavioral economics and decision analysis. Funding: This work was supported by the Vannevar Bush Faculty Fellowship and Multidisciplinary University Research Initiative [Grant W911NF-19-0217]. Supplemental Material: The online appendices are available at https://doi.org/10.1287/mnsc.2022.03683 .

A Behavioral Model of Rational Choice
Abstract. Introduction, 99. — I. Some general features of rational choice, 100.— II. The essential simplifications, 103. — III. Existence and uniqueness of

Can Consumers Make Affordable Care Affordable? The Value of Choice Architecture
Tens of millions of people are currently choosing health coverage on a state or federal health insurance exchange as part of the Patient Protection and Affordable Care Act. We examine how well people make these choices, how well they think they do, and what can be done to improve these choices. We conducted 6 experiments asking people to choose the most cost-effective policy using websites modeled on current exchanges. Our results suggest there is significant room for improvement. Without interventions, respondents perform at near chance levels and show a significant bias, overweighting out-of-pocket expenses and deductibles. Financial incentives do not improve performance, and decision-makers do not realize that they are performing poorly. However, performance can be improved quite markedly by providing calculation aids, and by choosing a “smart” default. Implementing these psychologically based principles could save purchasers of policies and taxpayers approximately 10 billion dollars every year.