







Bulletin of Mathematical Biology - Mathematical modelling is a widely used approach to understand and interpret clinical trial data. This modelling typically involves fitting mechanistic...
Evaluation framework for systems models
Abstract As decisions in drug development increasingly rely on predictions from mechanistic systems models, assessing the predictive capability of such models is becoming more important. Several frameworks for the development of quantitative systems pharmacology (QSP) models have been proposed. In this paper, we add to this body of work with a framework that focuses on the appropriate use of qualitative and quantitative model evaluation methods. We provide details and references for those wishing to apply these methods, which include sensitivity and identifiability analyses, as well as concepts such as validation and uncertainty quantification. Many of these methods have been used successfully in other fields, but are not as common in QSP modeling. We illustrate how to apply these methods to evaluate QSP models, and propose methods to use in two case studies. We also share examples of misleading results when inappropriate analyses are used.

Statistical Models Answer the Fundamental Clinical Question and Provide Clinical Trial Estimands – Statistical Thinking
Specific goals and estimation targets for randomized clinical trials have still not been well defined for general outcome variables. Proponents of causal inference calculus have claimed to define goals and estimands, but they have largely done so in a way that is not concordant with the most popular design, the parallel-group randomized trial. Causal inferential methods require the use of counterfactuals that are not informed by any data (outside of crossover studies) and make assumptions that are unverifiable, e.g., about the correlation structure of potential outcomes. Causal inferential structure also leads practitioners to act as if marginal treatment effect estimates are both helpful in decision making and transport to populations when in fact neither is true. Heterogeneity of participants within a treatment arm dictates heterogeneity of outcomes and heterogeneity of treatment effects when quantified on an absolute scale. Statistical models are best poised for estimation and causal inference that is specific to patient types. The increasing generality and robustness of statistical models bolsters the case. In this article I provide a succinct statement of the clinical goal of a parallel-group trial, and statistical estimands for it in the context of a general family of robust and efficient ordinal models that contain virtually all routinely used statistical models and tests as special cases.

Hierarchical Modeling
Hierarchical modeling is a powerful technique for modeling heterogeneity and, consequently, it is becoming increasingly ubiquitous in contemporary applied statistics. Unfortunately that ubiquitous application has not brought with it an equivalently ubiquitous understanding for how awkward these models can be to fit in practice. In this case study we dive deep into hierarchical models, from their theoretical motivations to their inherent degeneracies and the strategies needed to ensure robust computation. We'll learn not only how to use hierarchical models but also how to use them robustly.
A model qualification method for mechanistic physiological QSP models to support model‐informed drug development
Mechanistic physiological modeling is a scientific method that combines available data with scientific knowledge and engineering approaches to facilitate better understanding of biological systems, improve decision‐making, reduce risk, and increase efficiency in drug discovery and development. It is a type of quantitative systems pharmacology (QSP) approach that places drug‐specific properties in the context of disease biology. This tutorial provides a broadly applicable model qualification method (MQM) to ensure that mechanistic physiological models are fit for their intended purposes.

Recursive partitioning for heterogeneous causal effects
In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals for treatment effects, even with many covariates relative to the sample size, and without “sparsity” assumptions. We propose an “honest” approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation. Our approach builds on regression tree methods, modified to optimize for goodness of fit in treatment effects and to account for honest estimation. Our model selection criterion anticipates that bias will be eliminated by honest estimation and also accounts for the effect of making additional splits on the variance of treatment effect estimates within each subpopulation. We address the challenge that the “ground truth” for a causal effect is not observed for any individual unit, so that standard approaches to cross-validation must be modified. Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches. Honest estimation requires estimating the model with a smaller sample size; the cost in terms of mean squared error of treatment effects for our preferred method ranges between 7–22%.

Reducing Bias in Observational Studies Using Subclassification on the Propensity Score
The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Previous theoretical arguments have shown that subclassification on the propensity score will balance all observed covariates. Subclassification on an estimated propensity score is illustrated, using observational data on treatments for coronary artery disease. Five subclasses defined by the estimated propensity score are constructed that balance 74 covariates, and thereby provide estimates of treatment effects using direct adjustment. These subclasses are applied within subpopulations, and model-based adjustments are then used to provide estimates of treatment effects within these subpopulations. Two appendixes address theoretical issues related to the application: the effectiveness of subclassification on the propensity score in removing bias, and balancing properties of propensity scores with incomplete data.
Choosing informative priors in Bayesian regression models: a simulation study and tutorial using Stan and R
BackgroundBayesian regression models provide a robust framework for complex data analysis, which is particularly advantageous in scenarios with small sample sizes, common in psychology or medical research. However, specifying appropriate prior distributions that incorporate existing knowledge to regularize model parameters remains a challenge for many researchers. This can lead to unstable or implausible estimates. This study aims to demonstrate the impact of different prior distributions on regression models and to provide a practical guide for choosing and justifying informative priors to produce more stable and credible results.MethodsThe study involved two parts. First, a simulation study was conducted to systematically assess the sensitivity of Bayesian linear regression models to prior specification. We systematically varied the sample size, prior location, and prior scale to observe their impact on posterior estimates for a known true effect size. Second, a case–control study using real-world patient data (N = 526) demonstrated the practical application of choosing informative priors. Bayesian logistic regression models were used to analyze the relationship between severe dementia and fall incidence, comparing results from priors based on existing literature (“believer”), conservative priors (“agnostic”), and priors assuming an opposite effect (“skeptical”).ResultsThe simulation study showed that strongly informative priors had a substantial influence on posterior estimates, particularly for smaller sample sizes. As the sample size increased, the influence of the data increased, and the estimates converged toward the true effect. In the case–control study, a standard frequentist logistic regression produced an odds ratio of 8.87 with a very wide and unstable confidence interval (1.66–165.19), likely due to data sparsity. In contrast, a Bayesian model using a moderately informative “believer” prior derived from existing research yielded a more stable and plausible odds ratio of 4.01 with a substantially narrower credible interval (1.99–8.78).ConclusionCareful and transparent specification of informative priors is a critical tool in Bayesian analysis, especially when data are sparse. By incorporating justified evidence-based assumptions, researchers can regularize models to prevent implausible outcomes and produce more stable, interpretable, and credible results. This approach enhances the robustness of statistical inference in fields where small sample sizes are a frequent challenge.

mediation: R Package for Causal Mediation Analysis
In this paper, we describe the R package mediation for conducting causal mediation analysis in applied empirical research. In many scientific disciplines, the goal of researchers is not only estimating causal effects of a treatment but also understanding the process in which the treatment causally affects the outcome. Causal mediation analysis is frequently used to assess potential causal mechanisms. The mediation package implements a comprehensive suite of statistical tools for conducting such an analysis. The package is organized into two distinct approaches. Using the model-based approach, researchers can estimate causal mediation effects and conduct sensitivity analysis under the standard research design. Furthermore, the design-based approach provides several analysis tools that are applicable under different experimental designs. This approach requires weaker assumptions than the model-based approach. We also implement a statistical method for dealing with multiple (causally dependent) mediators, which are often encountered in practice. Finally, the package also offers a methodology for assessing causal mediation in the presence of treatment noncompliance, a common problem in randomized trials.

Automated Scale Reduction of Nonlinear <span style="font-variant:small-caps;">QSP</span> Models With an Illustrative Application to a Bone Biology System
Integrating quantitative systems pharmacology ( QSP ) into pharmacokinetics/pharmacodynamics ( PKPD ) has resulted in models that are highly complex and often not amenable to further exploration via estimation or design. Because QSP models are usually depicted using nonlinear differential equations it is not straightforward to apply some model reduction techniques, such as proper lumping. In this study, we explore the combined use of linearization and proper lumping as a general method to simplification of a nonlinear QSP model. We illustrate this with a bone biology model and the reduced model was then applied to describe bone mineral density ( BMD ) changes due to denosumab dosing. The methodologies used in this study can be applied to other multiscale models for developing a mechanism‐based structural model for future analyses.

Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies
Presents a discussion of matching, randomization, random sampling, and other methods of controlling extraneous variation. The objective was to specify the benefits of randomization in estimating causal effects of treatments. It is concluded that randomization should be employed whenever possible but that the use of carefully controlled nonrandomized data to estimate causal effects is a reasonable and necessary procedure in many cases.
Training Data Attribution via Approximate Unrolling
Many training data attribution (TDA) methods aim to estimate how a model's behavior would change if one or more data points were removed from the training set. Methods based on implicit differentiation, such as influence functions, can be made computationally efficient, but fail to account for underspecification, the implicit bias of the optimization algorithm, or multi-stage training pipelines. By contrast, methods based on unrolling address these issues but face scalability challenges. In this work, we connect the implicit-differentiation-based and unrolling-based approaches and combine their benefits by introducing Source, an approximate unrolling-based TDA method that is computed using an influence-function-like formula. While being computationally efficient compared to unrolling-based approaches, Source is suitable in cases where implicit-differentiation-based approaches struggle, such as in non-converged models and multi-stage training pipelines. Empirically, Source outperforms existing TDA techniques in counterfactual prediction, especially in settings where implicit-differentiation-based approaches fall short.
Machine Learning Who to Nudge: Causal vs Predictive Targeting in a Field Experiment on Student Financial Aid Renewal
In many settings, interventions may be more effective for some individuals than others, so that targeting interventions may be beneficial. We analyze the value of targeting in the context of a large-scale field experiment with over 53,000 college students, where the goal was to use "nudges" to encourage students to renew their financial-aid applications before a non-binding deadline. We begin with baseline approaches to targeting. First, we target based on a causal forest that estimates heterogeneous treatment effects and then assigns students to treatment according to those estimated to have the highest treatment effects. Next, we evaluate two alternative targeting policies, one targeting students with low predicted probability of renewing financial aid in the absence of the treatment, the other targeting those with high probability. The predicted baseline outcome is not the ideal criterion for targeting, nor is it a priori clear whether to prioritize low, high, or intermediate predicted probability. Nonetheless, targeting on low baseline outcomes is common in practice, for example because the relationship between individual characteristics and treatment effects is often difficult or impossible to estimate with historical data. We propose hybrid approaches that incorporate the strengths of both predictive approaches (accurate estimation) and causal approaches (correct criterion); we show that targeting intermediate baseline outcomes is most effective in our specific application, while targeting based on low baseline outcomes is detrimental. In one year of the experiment, nudging all students improved early filing by an average of 6.4 percentage points over a baseline average of 37% filing, and we estimate that targeting half of the students using our preferred policy attains around 75% of this benefit.

Introduction to Statistical Mediation Analysis
This volume introduces the statistical, methodological, and conceptual aspects of mediation analysis. Applications from health, social, and developmental psychology, sociology, communication, exercise science, and epidemiology are emphasized throughout. Single-mediator, multilevel, and longitudinal models are reviewed. The author's goal is to help the reader apply mediation analysis to their own data and understand its limitations. Each chapter features an overview, numerous worked examples, a summary, and exercises (with answers to the odd numbered questions). The accompanying CD contains outputs described in the book from SAS, SPSS, LISREL, EQS, MPLUS, and CALIS, and a program to simulate the model. The notation used is consistent with existing literature on mediation in psychology. The book opens with a review of the types of research questions the mediation model addresses. Part II describes the estimation of mediation effects including assumptions, statistical tests, and the construction of confidence limits. Advanced models including mediation in path analysis, longitudinal models, multilevel data, categorical variables, and mediation in the context of moderation are then described. The book closes with a discussion of the limits of mediation analysis, additional approaches to identifying mediating variables, and future directions. Introduction to Statistical Mediation Analysis is intended for researchers and advanced students in health, social, clinical, and developmental psychology as well as communication, public health, nursing, epidemiology, and sociology. Some exposure to a graduate level research methods or statistics course is assumed. The overview of mediation analysis and the guidelines for conducting a mediation analysis will be appreciated by all readers.

Assessing and mitigating batch effects in large-scale omics studies
Batch effects in omics data are notoriously common technical variations unrelated to study objectives, and may result in misleading outcomes if uncorrected, or hinder biomedical discovery if over-corrected. Assessing and mitigating batch effects is crucial for ensuring the reliability and reproducibility of omics data and minimizing the impact of technical variations on biological interpretation. In this review, we highlight the profound negative impact of batch effects and the urgent need to address this challenging problem in large-scale omics studies. We summarize potential sources of batch effects, current progress in evaluating and correcting them, and consortium efforts aiming to tackle them.

Randomized Controlled Trials without Data Retention
Amidst rising appreciation for privacy and data usage rights, researchers have increasingly acknowledged the principle of data minimization, which holds that the accessibility, collection, and retention of subjects' data should be kept to the bare amount needed to answer focused research questions. Applying this principle to randomized controlled trials (RCTs), this paper presents algorithms for making accurate inferences from RCTs under stringent data retention and anonymization policies. In particular, we show how to use recursive algorithms to construct running estimates of treatment effects in RCTs, which allow individualized records to be deleted or anonymized shortly after collection. Devoting special attention to non-i.i.d. data, we further show how to draw robust inferences from RCTs by combining recursive algorithms with bootstrap and federated strategies.

Can Revealed Preferences Clarify LLM Alignment and Steering?
LLMs are increasingly used to make or support high-stakes decisions under uncertainty, where alignment depends not only on factual accuracy but on how models weigh tradeoffs between different outcomes. We present an empirical pipeline for estimating the implied preferences that an LLM's observed choices optimize: we elicit the model's probability distribution over unknowns along with the choice it would make for the decision task and then fit a discrete choice model to recover the cost function that best rationalizes the model's decisions. We show how this revealed-preference description allows rigorous evaluation of whether models behave in a consistently goal-directed way, whether they can verbalize a description of their objectives which matches their revealed decision policy, and whether prompting can reliably steer those policies to implement a user-specified cost function. We apply this evaluation across four medical diagnosis domains and multiple frontier and open-source models. We find that while many models have a nontrivial degree of internal coherence, they also have significant weaknesses in faithfully reporting or adopting preferences in response to user direction.
