







RCTs do NOT estimate average effects; they also rarely EXPLAIN that which they have demonstrated.
Why Adjusted Regression Coefficients Are Less Descriptive Than They Look | R Psychologist
An interactive article about how regression adjustment works and the pitfalls of conditional associations in descriptive studies.

“Statistical Significance” and Statistical Reporting: Moving Beyond Binary
Null hypothesis significance testing (NHST) is the default approach to statistical analysis and reporting in marketing and the biomedical and social sciences more broadly. Despite its default role, NHST has long been criticized by both statisticians and applied researchers, including those within marketing. Therefore, the authors propose a major transition in statistical analysis and reporting. Specifically, they propose moving beyond binary: abandoning NHST as the default approach to statistical analysis and reporting. To facilitate this, they briefly review some of the principal problems associated with NHST. They next discuss some principles that they believe should underlie statistical analysis and reporting. They then use these principles to motivate some guidelines for statistical analysis and reporting. They next provide some examples that illustrate statistical analysis and reporting that adheres to their principles and guidelines. They conclude with a brief discussion.

Why Most Published Research Findings Are False
Summary There is increasing concern that most current published research findings are false. The probability that a research claim is true may depend on study power and bias, the number of other studies on the same question, and, importantly, the ratio of true to no relationships among the relationships probed in each scientific field. In this framework, a research finding is less likely to be true when the studies conducted in a field are smaller; when effect sizes are smaller; when there is a greater number and lesser preselection of tested relationships; where there is greater flexibility in designs, definitions, outcomes, and analytical modes; when there is greater financial and other interest and prejudice; and when more teams are involved in a scientific field in chase of statistical significance. Simulations show that for most study designs and settings, it is more likely for a research claim to be false than true. Moreover, for many current scientific fields, claimed research findings may often be simply accurate measures of the prevailing bias. In this essay, I discuss the implications of these problems for the conduct and interpretation of research.
On the Accuracy of Influence Functions for Measuring Group Effects
Influence functions estimate the effect of removing a training point on a model without the need to retrain. They are based on a first-order Taylor approximation that is guaranteed to be accurate for sufficiently small changes to the model, and so are commonly used to study the effect of individual points in large datasets. However, we often want to study the effects of large groups of training points, e.g., to diagnose batch effects or apportion credit between different data sources. Removing such large groups can result in significant changes to the model. Are influence functions still accurate in this setting? In this paper, we find that across many different types of groups and for a range of real-world datasets, the predicted effect (using influence functions) of a group correlates surprisingly well with its actual effect, even if the absolute and relative errors are large. Our theoretical analysis shows that such strong correlation arises only under certain settings and need not hold in general, indicating that real-world datasets have particular properties that allow the influence approximation to be accurate.
Assessing and mitigating batch effects in large-scale omics studies
Batch effects in omics data are notoriously common technical variations unrelated to study objectives, and may result in misleading outcomes if uncorrected, or hinder biomedical discovery if over-corrected. Assessing and mitigating batch effects is crucial for ensuring the reliability and reproducibility of omics data and minimizing the impact of technical variations on biological interpretation. In this review, we highlight the profound negative impact of batch effects and the urgent need to address this challenging problem in large-scale omics studies. We summarize potential sources of batch effects, current progress in evaluating and correcting them, and consortium efforts aiming to tackle them.

The natural selection of bad science
Abstract. Poor research design and data analysis encourage false-positive findings. Such poor methods persist despite perennial calls for improvement, sugg

Prof. Lee Cronin on Twitter / X
It is easy to see that biology does things that cannot be computed in advanced because the future is not controlled by statistics, it is controlled by creative actions. This means today’s statistically-driven GenAI is fundamentally unintelligent.— Prof. Lee Cronin (@leecronin) March 28, 2026
The C-Word: Scientific Euphemisms Do Not Improve Causal Inference From Observational Data
Causal inference is a core task of science. However, authors and editors often refrain from explicitly acknowledging the causal goal of research projects; they refer to causal effect estimates as associational estimates. This commentary argues that using the term “causal” is necessary to improve the quality of observational research. Specifically, being explicit about the causal objective of a study reduces ambiguity in the scientific question, errors in the data analysis, and excesses in the interpretation of the results.

Statistical Modeling, Causal Inference, and Social Science
I saw in a recent issue of the Times Literary Supplement that you have been critical of the “chambermaid” study which purported to show that people were losing weight without changing their diet or exercise. I agree that this study did not show what it claimed.
In Praise of Moderation: Suggestions for the Scope and Use of Pre-Analysis Plans for RCTs in Economics
Pre-Analysis Plans (PAPs) for randomized evaluations are becoming increasingly common in Economics, but their definition remains unclear and their practical applications therefore vary widely. Based on our collective experiences as researchers and editors, we articulate a set of principles for the ex-ante scope and ex-post use of PAPs. We argue that the key benefits of a PAP can usually be realized by completing the registration fields in the AEA RCT Registry. Specific cases where more detail may be warranted include when subgroup analysis is expected to be particularly important, or a party to the study has a vested interest. However, a strong norm for more detailed pre-specification can be detrimental to knowledge creation when implementing field experiments in the real world. An ex-post requirement of strict adherence to pre-specified plans, or the discounting of non-pre-specified work, may mean that some experiments do not take place, or that interesting observations and new theories are not explored and reported. Rather, we recommend that the final research paper be written and judged as a distinct object from the “results of the PAP”; to emphasize this distinction, researchers could consider producing a short, publicly available report (the “populated PAP”) that populates the PAP to the extent possible and briefly discusses any barriers to doing so.

Small Telescopes - Uri Simonsohn, 2015
This article introduces a new approach for evaluating replication results. It combines effect-size estimation with hypothesis testing, assessing the extent to w...

Andreas De Block on Twitter / X
It took Nature three years to publish this rebuttal, which clearly demonstrates the fundamental flaws in the original paper. In the meantime, the paper's conclusions influenced scientific debate and policy decisions. Such delays in correcting flawed research do real damage. https://t.co/QI3Ei3PE70— Andreas De Block (@DeblockBlock) August 13, 2026
99% impossible: A valid, or falsifiable, internal meta-analysis.
Rutger Bregman on Twitter / X
Devastating review of the degrowth literature (561 studies): --> 'few studies use quantitative or qualitative data...' --> [those that do] 'tend to include small samples or focus on non-representative cases' -->'large majority (almost 90%) are opinions rather than analysis' pic.twitter.com/1OQuUhzfk5— Rutger Bregman (@rcbregman) September 4, 2024

#SciComment Nice example of protecting ECRs from critiques that they had no power to avoid...