







Background Previous research has shown a correlation between the imposition of sanctions and worsening health conditions in target countries. However, the direction of causality in this relationship remains unclear. No study has yet examined the effects of sanctions on age-specific mortality rates in cross-country panel data using methods designed to address causal identification in observational data. Methods In this cross-national panel data analysis, we analysed the effect on health of sanctions using a panel dataset of age-specific mortality rates and sanctions episodes for 152 countries between 1971 and 2021. We apply a range of methods designed to address causal questions using observational data, including entropy balancing, Granger causality, event-study representations, and instrumental variables. Findings Our findings showed a significant causal association between sanctions and increased mortality. We found the strongest effects for unilateral, economic, and US sanctions, whereas we found no statistical evidence of an effect for UN sanctions. Mortality effects ranged from 8·4 log points (95% CI 3·9–13·0) for children younger than 5 years to 2·4 log points (0·9–4·0) for individuals aged 60–80 years. We estimated that unilateral sanctions were associated with an annual toll of 564 258 deaths (95% CI 367 838–760 677), similar to the global mortality burden associated with armed conflict. Interpretation Sanctions have substantial adverse effects on public health, with a death toll similar to that of wars. Our findings underscore the need to rethink sanctions as a foreign-policy tool, highlighting the importance of exercising restraint in their use and seriously considering efforts to reform their design. Funding The Center for Economic and Policy Research.
The central role of the propensity score in observational studies for causal effects
Abstract. The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Both large and

Machine Learning Who to Nudge: Causal vs Predictive Targeting in a Field Experiment on Student Financial Aid Renewal
In many settings, interventions may be more effective for some individuals than others, so that targeting interventions may be beneficial. We analyze the value of targeting in the context of a large-scale field experiment with over 53,000 college students, where the goal was to use "nudges" to encourage students to renew their financial-aid applications before a non-binding deadline. We begin with baseline approaches to targeting. First, we target based on a causal forest that estimates heterogeneous treatment effects and then assigns students to treatment according to those estimated to have the highest treatment effects. Next, we evaluate two alternative targeting policies, one targeting students with low predicted probability of renewing financial aid in the absence of the treatment, the other targeting those with high probability. The predicted baseline outcome is not the ideal criterion for targeting, nor is it a priori clear whether to prioritize low, high, or intermediate predicted probability. Nonetheless, targeting on low baseline outcomes is common in practice, for example because the relationship between individual characteristics and treatment effects is often difficult or impossible to estimate with historical data. We propose hybrid approaches that incorporate the strengths of both predictive approaches (accurate estimation) and causal approaches (correct criterion); we show that targeting intermediate baseline outcomes is most effective in our specific application, while targeting based on low baseline outcomes is detrimental. In one year of the experiment, nudging all students improved early filing by an average of 6.4 percentage points over a baseline average of 37% filing, and we estimate that targeting half of the students using our preferred policy attains around 75% of this benefit.

Why Do Americans Have Shorter Life Expectancy and Worse Health Than Do People in Other High-Income Countries?
Americans lead shorter and less healthy lives than do people in other high-income countries. We review the evidence and explanations for these variations in longevity and health. Our overview suggests that the US health disadvantage applies to multiple mortality and morbidity outcomes. The American health disadvantage begins at birth and extends across the life course, and it is particularly marked for American women and for regions in the US South and Midwest. Proposed explanations include differences in health care, individual behaviors, socioeconomic inequalities, and the built physical environment. Although these factors may contribute to poorer health in America, a focus on proximal causes fails to adequately account for the ubiquity of the US health disadvantage across the life course. We discuss the role of specific public policies and conclude that while multiple causes are implicated, crucial differences in social policy might underlie an important part of the US health disadvantage.

Global aid cuts could lead to 23 million deaths by 2030, study estimates
Lab tests are negative for more than 200 pathogens, but testing continues.

The C-Word: Scientific Euphemisms Do Not Improve Causal Inference From Observational Data
Causal inference is a core task of science. However, authors and editors often refrain from explicitly acknowledging the causal goal of research projects; they refer to causal effect estimates as associational estimates. This commentary argues that using the term “causal” is necessary to improve the quality of observational research. Specifically, being explicit about the causal objective of a study reduces ambiguity in the scientific question, errors in the data analysis, and excesses in the interpretation of the results.

Structural adjustment: damages, reparations and pathways to non-recurrence
Beginning in the 1980s and 1990s, the International Monetary Fund (IMF) and the World Bank implemented neoliberal structural adjustment programmes (SAPs) across most countries in Asia, Africa and Latin America. SAPs imposed austerity, privatisation and economic deregulation and have been associated with severe negative impacts on human welfare, including (a) declining real wages and working-class consumption, (b) increased rates of poverty and basic-needs deprivation, (c) increased neonatal and maternal mortality and (d) reduced health system access. Structural adjustment also created conditions for increased financial outflows and drain from the global South through unequal exchange. This paper reviews evidence of these damages and proposes possible options for reparations and distributive justice. We argue that the IMF and the World Bank should be democratised and restructured—or otherwise replaced by alternative institutions—to prevent further harm.
mediation: R Package for Causal Mediation Analysis
In this paper, we describe the R package mediation for conducting causal mediation analysis in applied empirical research. In many scientific disciplines, the goal of researchers is not only estimating causal effects of a treatment but also understanding the process in which the treatment causally affects the outcome. Causal mediation analysis is frequently used to assess potential causal mechanisms. The mediation package implements a comprehensive suite of statistical tools for conducting such an analysis. The package is organized into two distinct approaches. Using the model-based approach, researchers can estimate causal mediation effects and conduct sensitivity analysis under the standard research design. Furthermore, the design-based approach provides several analysis tools that are applicable under different experimental designs. This approach requires weaker assumptions than the model-based approach. We also implement a statistical method for dealing with multiple (causally dependent) mediators, which are often encountered in practice. Finally, the package also offers a methodology for assessing causal mediation in the presence of treatment noncompliance, a common problem in randomized trials.

Recursive partitioning for heterogeneous causal effects
In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals for treatment effects, even with many covariates relative to the sample size, and without “sparsity” assumptions. We propose an “honest” approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation. Our approach builds on regression tree methods, modified to optimize for goodness of fit in treatment effects and to account for honest estimation. Our model selection criterion anticipates that bias will be eliminated by honest estimation and also accounts for the effect of making additional splits on the variance of treatment effect estimates within each subpopulation. We address the challenge that the “ground truth” for a causal effect is not observed for any individual unit, so that standard approaches to cross-validation must be modified. Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches. Honest estimation requires estimating the model with a smaller sample size; the cost in terms of mean squared error of treatment effects for our preferred method ranges between 7–22%.

Integrative experiments identify how punishment affects welfare in public goods games
Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena. , Editor’s summary People face conflicts between maximizing personal gain versus supporting collective interests. If we cooperatively recycle or donate to charities, it benefits society, but it also costs us time and resources that could be selfishly preserved for ourselves. We impose penalties to deter those undesirable or selfish behaviors, but under what conditions do punishments or penalties effectively modify behavior to benefit group welfare? Alsobay et al . systematically and simultaneously varied 14 factors together instead of in isolation. Punishment was unequivocally most effective when paired with consistent communication, particularly over time. Another effective factor was “opting out” or withdrawing some, but not all, endowments already in the public fund. These methodological advances revealed when, rather than whether, punishment works. —Ekeoma Uzogara , INTRODUCTION Human societies face many situations where individual and collective interests conflict, often referred to as social dilemmas. Costly peer punishment has been studied for more than 25 years in public goods games (stylized behavioral experiments in which individuals decide how much to contribute to a shared pool that benefits everyone) as a mechanism to promote cooperation. Prior research has identified many contextual factors that moderate punishment’s effectiveness, including game length, communication, group size, punishment cost, and so on. However, the specific conditions under which punishment improves group welfare remain unclear. RATIONALE We argue that this lack of clarity derives from the dominant experimental paradigm, in which any given study manipulates only one or a few theoretically informed factors. Because such studies differ in many ways (different experimental procedures, populations), their results are often difficult to compare or integrate. Consequently, one can list many factors that have some effect, but cannot say how much each matters relative to the others, or how they work together, and as a result, cannot predict when punishment will help or harm welfare in new settings. To address this fundamental knowledge gap, we use an integrative experimental design and systematically vary 14 parameters across 360 conditions (147,618 decisions from 7100 participants) to elucidate when punishment improves versus undermines welfare in public goods games, which factors matter most, and how they interact. RESULTS The effect of punishment on welfare ranged from 43% improvement to 44% reduction depending on the specific combination of game parameters. To characterize this heterogeneity, we trained a model that outperformed all 553 human forecasters (laypeople and experts) in predicting whether punishment would help or harm welfare in new experiments. Communication emerged as roughly three times more important than any other factor, followed by contribution framing (opt in versus opt out), contribution type (variable versus all-or-nothing), game length, and peer outcome visibility (whether participants can see others’ earnings). These factors often interact. For example, longer games enhance punishment’s effectiveness only when communication is available, and contribution framing effects depend on both contribution type and outcome visibility. CONCLUSION Many phenomena in social science are shaped by many factors whose interactions are consequential, yet the dominant experimental paradigm often limits its inquiry to “does a given effect exist?” and examines hypothesized factors in isolation. As a result, research programs can accumulate many partial explanations without a clear picture of how they combine to determine outcomes across settings. Knowing that factors matter individually is fundamentally different from knowing how much each matters and how they interact. The integrative approach implemented here offers one way forward. It varies many factors simultaneously within a shared design space, evaluates models by their predictive accuracy on new experiments, and probes those models to constrain and develop theory. Our hope is that integrative experiment designs, combined with models that integrate prediction and explanation, represent a path toward more cumulative social science. Integrative experiment reveals when punishment helps versus harms. We systematically varied 14 design parameters across 360 experimental conditions. The effect of punishment on cooperation efficiency ranged from −44% to +43% depending on the specific game parameters. Communication emerged as three times more important than any other factor, followed by contribution framing, contribution type, and game length.

Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies
Presents a discussion of matching, randomization, random sampling, and other methods of controlling extraneous variation. The objective was to specify the benefits of randomization in estimating causal effects of treatments. It is concluded that randomization should be employed whenever possible but that the use of carefully controlled nonrandomized data to estimate causal effects is a reasonable and necessary procedure in many cases.
A Statistical Interrogation of “The Case for Causality, Part 1” by Rausch and Haidt – Matthew B. Jané
I don’t know anything about the literature on social media and mental health so my focus on this post is to interrogate the statistical approach taken by the article written by Zach Rausch and Jonathon Haidt (link here) and to some extent the original meta-analysis by Ferguson.

Mechanism Experiments and Policy Evaluations
Randomized controlled trials are increasingly used to evaluate policies. How can we make these experiments as useful as possible for policy purposes? We argue greater use should be made of experiments that identify the behavioral mechanisms that are central to clearly specified policy questions, what we call "mechanism experiments." These types of experiments can be of great policy value even if the intervention that is tested (or its setting) does not correspond exactly to any realistic policy option.
Statistical Models Answer the Fundamental Clinical Question and Provide Clinical Trial Estimands – Statistical Thinking
Specific goals and estimation targets for randomized clinical trials have still not been well defined for general outcome variables. Proponents of causal inference calculus have claimed to define goals and estimands, but they have largely done so in a way that is not concordant with the most popular design, the parallel-group randomized trial. Causal inferential methods require the use of counterfactuals that are not informed by any data (outside of crossover studies) and make assumptions that are unverifiable, e.g., about the correlation structure of potential outcomes. Causal inferential structure also leads practitioners to act as if marginal treatment effect estimates are both helpful in decision making and transport to populations when in fact neither is true. Heterogeneity of participants within a treatment arm dictates heterogeneity of outcomes and heterogeneity of treatment effects when quantified on an absolute scale. Statistical models are best poised for estimation and causal inference that is specific to patient types. The increasing generality and robustness of statistical models bolsters the case. In this article I provide a succinct statement of the clinical goal of a parallel-group trial, and statistical estimands for it in the context of a general family of robust and efficient ordinal models that contain virtually all routinely used statistical models and tests as special cases.

OpenSanctions: Supreme Data on Supreme Leaders
The open-source database of sanctions, watchlists, and politically exposed persons — aggregating hundreds of sources and relied on by compliance teams, investigators, and journalists.

Reducing Bias in Observational Studies Using Subclassification on the Propensity Score
The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Previous theoretical arguments have shown that subclassification on the propensity score will balance all observed covariates. Subclassification on an estimated propensity score is illustrated, using observational data on treatments for coronary artery disease. Five subclasses defined by the estimated propensity score are constructed that balance 74 covariates, and thereby provide estimates of treatment effects using direct adjustment. These subclasses are applied within subpopulations, and model-based adjustments are then used to provide estimates of treatment effects within these subpopulations. Two appendixes address theoretical issues related to the application: the effectiveness of subclassification on the propensity score in removing bias, and balancing properties of propensity scores with incomplete data.
Adult age differences in monetary decisions with real and hypothetical reward
Abstract Age differences in monetary decisions may emerge because younger and older adults perceive the value of outcomes differently. Yet, age‐differential effects of monetary rewards on decisions are not well understood. Most laboratory studies on aging and decision making have used scenarios in which rewards were merely hypothetical (decisions did not have any real consequences) or in which only small amounts of money were at stake. In the current study, we compared younger adults' (20–29 years) and older adults' (61–82 years) decisions in probabilistic choice problems with real or hypothetical rewards. Decision‐contingent rewards were in a typical range of previous studies (gains of up to ~4.25 USD) or substantially scaled up (gains of up to ~85 USD per participant). Reward type (real vs. hypothetical) affected decision quality, including value maximization, switching between options, and dominance violations (choices of an option that was inferior to another option in all respects). Decision quality was markedly better with real than hypothetical rewards in older adults and correlated with numeracy in both age groups. However, we found no evidence that reward type affected people's risk preferences. Overall, the findings portray a fairly positive picture regarding the use of hypothetical scenarios to assess preferences: With carefully prepared instructions, people from different age groups indicate preferences in hypothetical scenarios that match their decisions with real and much higher rewards. One advantage of using real rewards is that they help to reduce decision noise.
