







The central role of the propensity score in observational studies for causal effects
Abstract. The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Both large and

Statistical Models Answer the Fundamental Clinical Question and Provide Clinical Trial Estimands – Statistical Thinking
Specific goals and estimation targets for randomized clinical trials have still not been well defined for general outcome variables. Proponents of causal inference calculus have claimed to define goals and estimands, but they have largely done so in a way that is not concordant with the most popular design, the parallel-group randomized trial. Causal inferential methods require the use of counterfactuals that are not informed by any data (outside of crossover studies) and make assumptions that are unverifiable, e.g., about the correlation structure of potential outcomes. Causal inferential structure also leads practitioners to act as if marginal treatment effect estimates are both helpful in decision making and transport to populations when in fact neither is true. Heterogeneity of participants within a treatment arm dictates heterogeneity of outcomes and heterogeneity of treatment effects when quantified on an absolute scale. Statistical models are best poised for estimation and causal inference that is specific to patient types. The increasing generality and robustness of statistical models bolsters the case. In this article I provide a succinct statement of the clinical goal of a parallel-group trial, and statistical estimands for it in the context of a general family of robust and efficient ordinal models that contain virtually all routinely used statistical models and tests as special cases.

Recursive partitioning for heterogeneous causal effects
In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals for treatment effects, even with many covariates relative to the sample size, and without “sparsity” assumptions. We propose an “honest” approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation. Our approach builds on regression tree methods, modified to optimize for goodness of fit in treatment effects and to account for honest estimation. Our model selection criterion anticipates that bias will be eliminated by honest estimation and also accounts for the effect of making additional splits on the variance of treatment effect estimates within each subpopulation. We address the challenge that the “ground truth” for a causal effect is not observed for any individual unit, so that standard approaches to cross-validation must be modified. Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches. Honest estimation requires estimating the model with a smaller sample size; the cost in terms of mean squared error of treatment effects for our preferred method ranges between 7–22%.

Machine Learning Who to Nudge: Causal vs Predictive Targeting in a Field Experiment on Student Financial Aid Renewal
In many settings, interventions may be more effective for some individuals than others, so that targeting interventions may be beneficial. We analyze the value of targeting in the context of a large-scale field experiment with over 53,000 college students, where the goal was to use "nudges" to encourage students to renew their financial-aid applications before a non-binding deadline. We begin with baseline approaches to targeting. First, we target based on a causal forest that estimates heterogeneous treatment effects and then assigns students to treatment according to those estimated to have the highest treatment effects. Next, we evaluate two alternative targeting policies, one targeting students with low predicted probability of renewing financial aid in the absence of the treatment, the other targeting those with high probability. The predicted baseline outcome is not the ideal criterion for targeting, nor is it a priori clear whether to prioritize low, high, or intermediate predicted probability. Nonetheless, targeting on low baseline outcomes is common in practice, for example because the relationship between individual characteristics and treatment effects is often difficult or impossible to estimate with historical data. We propose hybrid approaches that incorporate the strengths of both predictive approaches (accurate estimation) and causal approaches (correct criterion); we show that targeting intermediate baseline outcomes is most effective in our specific application, while targeting based on low baseline outcomes is detrimental. In one year of the experiment, nudging all students improved early filing by an average of 6.4 percentage points over a baseline average of 37% filing, and we estimate that targeting half of the students using our preferred policy attains around 75% of this benefit.

Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies
Presents a discussion of matching, randomization, random sampling, and other methods of controlling extraneous variation. The objective was to specify the benefits of randomization in estimating causal effects of treatments. It is concluded that randomization should be employed whenever possible but that the use of carefully controlled nonrandomized data to estimate causal effects is a reasonable and necessary procedure in many cases.
Certifying and Removing Disparate Impact
What does it mean for an algorithm to be biased? In U.S. law, unintentional bias is encoded via disparate impact, which occurs when a selection process has widely different outcomes for different groups, even as it appears to be neutral. This legal determination hinges on a definition of a protected class (ethnicity, gender) and an explicit description of the process.When computers are involved, determining disparate impact (and hence bias) is harder. It might not be possible to disclose the process. In addition, even if the process is open, it might be hard to elucidate in a legal setting how the algorithm makes its decisions. Instead of requiring access to the process, we propose making inferences based on the data it uses.We present four contributions. First, we link disparate impact to a measure of classification accuracy that while known, has received relatively little attention. Second, we propose a test for disparate impact based on how well the protected class can be predicted from the other attributes. Third, we describe methods by which data might be made unbiased. Finally, we present empirical evidence supporting the effectiveness of our test for disparate impact and our approach for both masking bias and preserving relevant information in the data. Interestingly, our approach resembles some actual selection practices that have recently received legal scrutiny.

A Nonparametric Approach to Practical Identifiability of Nonlinear Mixed Effects Models
Mathematical modelling is a widely used approach to understand and interpret clinical trial data. This modelling typically involves fitting mechanistic mathematical models to data from individual trial participants. Despite the widespread adoption of this individual-based fitting, it is becoming increasingly common to take a hierarchical approach to parameter estimation, where modellers characterize the population parameter distributions, rather than considering each individual independently. This hierarchical parameter estimation is standard in pharmacometric modelling. However, many of the existing techniques for parameter identifiability do not immediately translate from the individual-based fitting to the hierarchical setting. In this work, we propose a nonparametric approach to study practical identifiability within a hierarchical parameter estimation framework. We focus on the commonly used nonlinear mixed effects framework and investigate two well-studied examples from the pharmacometrics and viral dynamics literature to illustrate the potential utility of our approach.

Why Adjusted Regression Coefficients Are Less Descriptive Than They Look | R Psychologist
An interactive article about how regression adjustment works and the pitfalls of conditional associations in descriptive studies.


Choice Bracketing
When making many choices, a person can broadly bracket them by assessing the consequences of all of them taken together, or narrowly bracket them by making each choice in isolation. We integrate research conducted in a wide range of decision contexts which shows that choice bracketing is an important determinant of behavior. Because broad bracketing allows people to take into account all the consequences of their actions, it generally leads to choices that yield higher utility. The evidence that we review, however, shows that people often fail to bracket broadly when it would be feasible for them to do so. In addition to documenting the diverse effects of bracketing, we also discuss factors that determine whether people bracket narrowly or broadly. We conclude with a discussion of normative aspects of bracketing and argue that there are some situations in which narrower bracketing results in superior decision making.

Beware of samples! A cognitive-ecological sampling approach to judgment biases.
Characterizing Fairness Over the Set of Good Models Under Selective Labels
Algorithmic risk assessments are used to inform decisions in a wide variety of high-stakes settings. Often multiple predictive models deliver similar overall performance but differ markedly in their predictions for individual cases, an empirical phenomenon known as the “Rashomon Effect.” These models may have different properties over various groups, and therefore have different predictive fairness properties. We develop a framework for characterizing predictive fairness properties over the set of models that deliver similar overall performance, or “the set of good models.” Our framework addresses the empirically relevant challenge of selectively labelled data in the setting where the selection decision and outcome are unconfounded given the observed data features. Our framework can be used to 1) audit for predictive bias; or 2) replace an existing model with one that has better fairness properties. We illustrate these use cases on a recidivism prediction task and a real-world credit-scoring task.
Choosing informative priors in Bayesian regression models: a simulation study and tutorial using Stan and R
BackgroundBayesian regression models provide a robust framework for complex data analysis, which is particularly advantageous in scenarios with small sample sizes, common in psychology or medical research. However, specifying appropriate prior distributions that incorporate existing knowledge to regularize model parameters remains a challenge for many researchers. This can lead to unstable or implausible estimates. This study aims to demonstrate the impact of different prior distributions on regression models and to provide a practical guide for choosing and justifying informative priors to produce more stable and credible results.MethodsThe study involved two parts. First, a simulation study was conducted to systematically assess the sensitivity of Bayesian linear regression models to prior specification. We systematically varied the sample size, prior location, and prior scale to observe their impact on posterior estimates for a known true effect size. Second, a case–control study using real-world patient data (N = 526) demonstrated the practical application of choosing informative priors. Bayesian logistic regression models were used to analyze the relationship between severe dementia and fall incidence, comparing results from priors based on existing literature (“believer”), conservative priors (“agnostic”), and priors assuming an opposite effect (“skeptical”).ResultsThe simulation study showed that strongly informative priors had a substantial influence on posterior estimates, particularly for smaller sample sizes. As the sample size increased, the influence of the data increased, and the estimates converged toward the true effect. In the case–control study, a standard frequentist logistic regression produced an odds ratio of 8.87 with a very wide and unstable confidence interval (1.66–165.19), likely due to data sparsity. In contrast, a Bayesian model using a moderately informative “believer” prior derived from existing research yielded a more stable and plausible odds ratio of 4.01 with a substantially narrower credible interval (1.99–8.78).ConclusionCareful and transparent specification of informative priors is a critical tool in Bayesian analysis, especially when data are sparse. By incorporating justified evidence-based assumptions, researchers can regularize models to prevent implausible outcomes and produce more stable, interpretable, and credible results. This approach enhances the robustness of statistical inference in fields where small sample sizes are a frequent challenge.

Associations of Sociodemographic and Neighborhood Vulnerability With Cardiovascular Health in Midlife Women in the United States
BACKGROUND: Many women have suboptimal cardiovascular health (CVH), which declines during midlife. Few studies have characterized CVH across the menopausal transition or identified its sociodemographic and neighborhood determinants. METHODS: We analyzed a prospective cohort of women in eastern Massachusetts enrolled during pregnancy (1999–2002) and followed to midlife (2019–2024). Exposures included household income, education, race and ethnicity, and neighborhood Social Vulnerability Index (categorized from very low [<20th percentile] to very high [≥80th percentile]; higher categories=greater neighborhood vulnerability). Women self-reported their menopause status using questionnaires. Using Life’s Essential 8, we derived CVH scores (0–100 points; higher score=better CVH) at 3-, 8-, 13-, 18-, and 23-year follow-up visits. Linear spline mixed-effect models examined associations of sociodemographics and neighborhood Social Vulnerability Index with differences in CVH across different menopause stages (premenopause, perimenopause, and postmenopause). RESULTS: Among 1200 women (mean enrollment age, 32.1 years; 67.5% Non-Hispanic White), 15.4% had household incomes ≤$40 000/y, 8.8% had ≤high school education, and 17.4% resided in very high Social Vulnerability Index neighborhoods. After covariate adjustment, women with lower income, lower education, or identifying as Non-Hispanic Black exhibited lower CVH across follow-up. Independent of individual sociodemographics, continued residence in vulnerable neighborhoods over time was associated with lower CVH and unfavorable CVH trajectories across follow-up. For example, residence in very high (versus very low) Social Vulnerability Index neighborhoods from enrollment to 3-year follow-up corresponded to mean CVH differences of −6.7 (95% CI, −12.3 to −1.2) at 3-year, −9.8 (95% CI, −15.6 to −4.0) at 8-year, −8.9 (95% CI, −13.9 to −3.9) at 13-year, −6.7 (95% CI, −12.9 to −0.5) at 18-year, and −7.2 (95% CI, −12.5 to −1.9) at 23-year follow-up, and with faster CVH score decline during premenopause (−0.62 points/y; 95% CI, −1.22 to −0.02). CONCLUSIONS: Women from disadvantaged sociodemographic backgrounds or residing in vulnerable neighborhoods exhibit poorer CVH across the menopausal transition, highlighting opportunities to optimize long-term CVH and mitigate cardiovascular disease risk.

Effects of international sanctions on age-specific mortality: a cross-national panel data analysis
Background Previous research has shown a correlation between the imposition of sanctions and worsening health conditions in target countries. However, the direction of causality in this relationship remains unclear. No study has yet examined the effects of sanctions on age-specific mortality rates in cross-country panel data using methods designed to address causal identification in observational data. Methods In this cross-national panel data analysis, we analysed the effect on health of sanctions using a panel dataset of age-specific mortality rates and sanctions episodes for 152 countries between 1971 and 2021. We apply a range of methods designed to address causal questions using observational data, including entropy balancing, Granger causality, event-study representations, and instrumental variables. Findings Our findings showed a significant causal association between sanctions and increased mortality. We found the strongest effects for unilateral, economic, and US sanctions, whereas we found no statistical evidence of an effect for UN sanctions. Mortality effects ranged from 8·4 log points (95% CI 3·9–13·0) for children younger than 5 years to 2·4 log points (0·9–4·0) for individuals aged 60–80 years. We estimated that unilateral sanctions were associated with an annual toll of 564 258 deaths (95% CI 367 838–760 677), similar to the global mortality burden associated with armed conflict. Interpretation Sanctions have substantial adverse effects on public health, with a death toll similar to that of wars. Our findings underscore the need to rethink sanctions as a foreign-policy tool, highlighting the importance of exercising restraint in their use and seriously considering efforts to reform their design. Funding The Center for Economic and Policy Research.
The persistence of cognitive biases in financial decisions across economic groups
While economic inequality continues to rise within countries, efforts to address it have been largely ineffective, particularly those involving behavioral approaches. It is often implied but not tested that choice patterns among low-income individuals may be a factor impeding behavioral interventions aimed at improving upward economic mobility. To test this, we assessed rates of ten cognitive biases across nearly 5000 participants from 27 countries. Our analyses were primarily focused on 1458 individuals that were either low-income adults or individuals who grew up in disadvantaged households but had above-average financial well-being as adults, known as positive deviants. Using discrete and complex models, we find evidence of no differences within or between groups or countries. We therefore conclude that choices impeded by cognitive biases alone cannot explain why some individuals do not experience upward economic mobility. Policies must combine both behavioral and structural interventions to improve financial well-being across populations.

The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Previous theoretical arguments have shown that subclassification on the propensity score will balance all observed covariates. Subclassification on an estimated propensity score is illustrated, using observational data on treatments for coronary artery disease. Five subclasses defined by the estimated propensity score are constructed that balance 74 covariates, and thereby provide estimates of treatment effects using direct adjustment. These subclasses are applied within subpopulations, and model-based adjustments are then used to provide estimates of treatment effects within these subpopulations. Two appendixes address theoretical issues related to the application: the effectiveness of subclassification on the propensity score in removing bias, and balancing properties of propensity scores with incomplete data.