







The view that the returns to educational investments are highest for early childhood interventions is widely held and stems primarily from several influential randomized trials—Abecedarian, Perry,...
A 2024 Review of Child Care and Early Learning in the United States
Updated data on child care and early learning in the United States illustrate the urgent need for holistic public policymaking and robust investments that support young children, families, and early educators.

Efficacy of infant simulator programmes to prevent teenage pregnancy: a school-based cluster randomised controlled trial in Western Australia
The infant simulator-based VIP programme did not achieve its aim of reducing teenage pregnancy. Girls in the intervention group were more likely to experience a birth or an induced abortion than those in the control group before they reached 20 years of age.

Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies
Presents a discussion of matching, randomization, random sampling, and other methods of controlling extraneous variation. The objective was to specify the benefits of randomization in estimating causal effects of treatments. It is concluded that randomization should be employed whenever possible but that the use of carefully controlled nonrandomized data to estimate causal effects is a reasonable and necessary procedure in many cases.
McMartin preschool trial
The McMartin preschool trial was a day care sexual abuse case in the 1980s, prosecuted by the Los Angeles District Attorney, Ira Reiner. Members of the McMartin family, who operated a preschool in Manhattan Beach, California, were charged with hundreds of acts of sexual abuse of children in their care. Accusations were made in 1983, with arrests and the pretrial investigation taking place from 1984 to 1987, and trials running from 1987 to 1990. The case lasted seven years but resulted in no convictions, and all charges were dropped in 1990. By the case's end, it had become the longest and most expensive series of criminal trials in American history.
Machine Learning Who to Nudge: Causal vs Predictive Targeting in a Field Experiment on Student Financial Aid Renewal
In many settings, interventions may be more effective for some individuals than others, so that targeting interventions may be beneficial. We analyze the value of targeting in the context of a large-scale field experiment with over 53,000 college students, where the goal was to use "nudges" to encourage students to renew their financial-aid applications before a non-binding deadline. We begin with baseline approaches to targeting. First, we target based on a causal forest that estimates heterogeneous treatment effects and then assigns students to treatment according to those estimated to have the highest treatment effects. Next, we evaluate two alternative targeting policies, one targeting students with low predicted probability of renewing financial aid in the absence of the treatment, the other targeting those with high probability. The predicted baseline outcome is not the ideal criterion for targeting, nor is it a priori clear whether to prioritize low, high, or intermediate predicted probability. Nonetheless, targeting on low baseline outcomes is common in practice, for example because the relationship between individual characteristics and treatment effects is often difficult or impossible to estimate with historical data. We propose hybrid approaches that incorporate the strengths of both predictive approaches (accurate estimation) and causal approaches (correct criterion); we show that targeting intermediate baseline outcomes is most effective in our specific application, while targeting based on low baseline outcomes is detrimental. In one year of the experiment, nudging all students improved early filing by an average of 6.4 percentage points over a baseline average of 37% filing, and we estimate that targeting half of the students using our preferred policy attains around 75% of this benefit.

Statistical Models Answer the Fundamental Clinical Question and Provide Clinical Trial Estimands – Statistical Thinking
Specific goals and estimation targets for randomized clinical trials have still not been well defined for general outcome variables. Proponents of causal inference calculus have claimed to define goals and estimands, but they have largely done so in a way that is not concordant with the most popular design, the parallel-group randomized trial. Causal inferential methods require the use of counterfactuals that are not informed by any data (outside of crossover studies) and make assumptions that are unverifiable, e.g., about the correlation structure of potential outcomes. Causal inferential structure also leads practitioners to act as if marginal treatment effect estimates are both helpful in decision making and transport to populations when in fact neither is true. Heterogeneity of participants within a treatment arm dictates heterogeneity of outcomes and heterogeneity of treatment effects when quantified on an absolute scale. Statistical models are best poised for estimation and causal inference that is specific to patient types. The increasing generality and robustness of statistical models bolsters the case. In this article I provide a succinct statement of the clinical goal of a parallel-group trial, and statistical estimands for it in the context of a general family of robust and efficient ordinal models that contain virtually all routinely used statistical models and tests as special cases.

Reducing Bias in Observational Studies Using Subclassification on the Propensity Score
The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Previous theoretical arguments have shown that subclassification on the propensity score will balance all observed covariates. Subclassification on an estimated propensity score is illustrated, using observational data on treatments for coronary artery disease. Five subclasses defined by the estimated propensity score are constructed that balance 74 covariates, and thereby provide estimates of treatment effects using direct adjustment. These subclasses are applied within subpopulations, and model-based adjustments are then used to provide estimates of treatment effects within these subpopulations. Two appendixes address theoretical issues related to the application: the effectiveness of subclassification on the propensity score in removing bias, and balancing properties of propensity scores with incomplete data.
Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?
A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while...

Children’s Sense of Fairness as Equal Respect
One influential view holds that children’s sense of fairness emerges at age 8 and is rooted in the development of an aversion to unequal resource distributions. Here, we suggest two amendments to this view. First, we argue and present evidence that children’s sense of fairness emerges already at age 3 in (and only in) the context of collaborative activities. This is because, in our theoretical view, collaboration creates a sense of equal respect among partners. Second, we argue and present evidence that children’s judgments about what is fair are essentially judgments about the social meaning of the distributive act; for example, children accept unequal distributions if the procedure gave everyone an equal chance (so-called distributive justice). Children thus respond to unequal (and other) distributions not based on material concerns, but rather based on interpersonal concerns: they want equal respect.
How Field Experiments in Economics Can Complement Psychological Research on Judgment Biases
This review summarizes results of field experiments examining individual behaviors across several market settings—from open-air markets to rideshare markets to tax-compliance markets—where people sort themselves into market roles wherein they make consequential decisions. Using three distinct examples from my own research on the endowment effect, left-digit bias, and omission bias, I showcase how field experiments can help researchers understand mediators, heterogeneity, and causal moderation involved in judgment biases in the field. In this manner, the review highlights that economic field experiments can serve an invaluable intellectual role alongside traditional laboratory research.

Higher Wages for Early Care and Education Workers Builds a Stronger System
Higher wages for early care and education workers in California is critical to expanding affordable child care.

How Are SNAP Benefits Spent? Evidence from a Retail Panel
We use a novel retail panel with detailed transaction records to study the effect of the Supplemental Nutrition Assistance Program (SNAP) on house-hold spending. We use administrative data to motivate three approaches to causal inference. The marginal propensity to consume SNAP-eligible food (MPCF) out of SNAP benefits is 0.5 to 0.6. The MPCF out of cash is much smaller. These patterns obtain even for households for whom SNAP benefits are economically equivalent to cash because their benefits are below their food spending. Using a semiparametric framework, we reject the hypothesis that households respect the fungibility of money. A model with mental accounting can match the facts.
Addressing Moderated Mediation Hypotheses: Theory, Methods, and Prescriptions
This article provides researchers with a guide to properly construe and conduct analyses of conditional indirect effects, commonly known as moderated mediation effects. We disentangle conflicting d...

Mechanism Experiments and Policy Evaluations
Randomized controlled trials are increasingly used to evaluate policies. How can we make these experiments as useful as possible for policy purposes? We argue greater use should be made of experiments that identify the behavioral mechanisms that are central to clearly specified policy questions, what we call "mechanism experiments." These types of experiments can be of great policy value even if the intervention that is tested (or its setting) does not correspond exactly to any realistic policy option.
Does AI stop children from learning?
New data show the peril and promise of the technology
