







Pre-Analysis Plans (PAPs) for randomized evaluations are becoming increasingly common in Economics, but their definition remains unclear and their practical applications therefore vary widely. Based on our collective experiences as researchers and editors, we articulate a set of principles for the ex-ante scope and ex-post use of PAPs. We argue that the key benefits of a PAP can usually be realized by completing the registration fields in the AEA RCT Registry. Specific cases where more detail may be warranted include when subgroup analysis is expected to be particularly important, or a party to the study has a vested interest. However, a strong norm for more detailed pre-specification can be detrimental to knowledge creation when implementing field experiments in the real world. An ex-post requirement of strict adherence to pre-specified plans, or the discounting of non-pre-specified work, may mean that some experiments do not take place, or that interesting observations and new theories are not explored and reported. Rather, we recommend that the final research paper be written and judged as a distinct object from the “results of the PAP”; to emphasize this distinction, researchers could consider producing a short, publicly available report (the “populated PAP”) that populates the PAP to the extent possible and briefly discusses any barriers to doing so.
Promises and Perils of Pre-analysis Plans
The purpose of this paper is to help think through the advantages and costs of rigorous pre-specification of statistical analysis plans in economics. A pre-analysis plan pre-specifies in a precise way the analysis to be run before examining the data. A researcher can specify variables, data cleaning procedures, regression specifications, and so on. If the regressions are pre-specified in advance and researchers are required to report all the results they pre-specify, data-mining problems are greatly reduced. I begin by laying out the basics of what a statistical analysis plan actually contains so those researchers unfamiliar with it can better understand how it is done. In so doing, I have drawn both on standards used in clinical trials, which are clearly specified by the Food and Drug Administration, as well as my own practical experience from writing these plans in economics contexts. I then lay out some of the advantages of pre-specified analysis plans, both for the scientific community as a whole and also for the researcher. I also explore some of the limitations and costs of such plans. I then review a few pieces of evidence that suggest that, in many contexts, the benefits of using pre-specified analysis plans may not be as high as one might have expected initially. For the most part, I will focus on the relatively narrow issue of pre-analysis for randomized controlled trials.
Against theory-motivated experimentation: Can random experimental choice lead to better theories?
Scientists must choose which among many experiments to perform. We study the epistemic success of experimental choice strategies proposed by philosophers of science or executed by scientists themselves. We develop a multi-agent model of the scientific process that jointly formalizes its core aspects: active experimentation, theorizing, and social learning. We find that agents who choose new experiments at random develop the most informative and predictive theories of the world. The agents aiming to confirm, falsify theories, or resolve theoretical disagreements end up with an illusion of epistemic success: they develop promising accounts for the data they collected, while misrepresenting the ground truth that they intended to learn about. Agents experimenting in these theory-motivated ways acquire less diverse or less representative samples from the ground truth that also turn out to be easier to account for. Random data collection, on the other hand, combines virtues of diverse and representative sampling from a target scientific domain which enables cumulative development of the successful theoretical accounts of it. We suggest that randomization, already a gold standard within experiments, is also beneficial at the level of experiments themselves.

Why Most Published Research Findings Are False
Summary There is increasing concern that most current published research findings are false. The probability that a research claim is true may depend on study power and bias, the number of other studies on the same question, and, importantly, the ratio of true to no relationships among the relationships probed in each scientific field. In this framework, a research finding is less likely to be true when the studies conducted in a field are smaller; when effect sizes are smaller; when there is a greater number and lesser preselection of tested relationships; where there is greater flexibility in designs, definitions, outcomes, and analytical modes; when there is greater financial and other interest and prejudice; and when more teams are involved in a scientific field in chase of statistical significance. Simulations show that for most study designs and settings, it is more likely for a research claim to be false than true. Moreover, for many current scientific fields, claimed research findings may often be simply accurate measures of the prevailing bias. In this essay, I discuss the implications of these problems for the conduct and interpretation of research.

Mechanism Experiments and Policy Evaluations
Randomized controlled trials are increasingly used to evaluate policies. How can we make these experiments as useful as possible for policy purposes? We argue greater use should be made of experiments that identify the behavioral mechanisms that are central to clearly specified policy questions, what we call "mechanism experiments." These types of experiments can be of great policy value even if the intervention that is tested (or its setting) does not correspond exactly to any realistic policy option.
Designing Information Provision Experiments
Information provision experiments allow researchers to test economic theories and answer policy-relevant questions by varying the information set available to respondents. We survey the emerging literature using information provision experiments in economics and discuss applications in macroeconomics, finance, political economy, public economics, labor economics, and health economics. We also discuss design considerations and provide best-practice recommendations on how to (i) measure beliefs; (ii) design the information intervention; (iii) measure belief updating; (iv) deal with potential confounds, such as experimenter demand effects; and (v) recruit respondents using online panels. We finally discuss typical effect sizes and provide sample size recommendations.
Research topic choice: Motivations, strategies, and consequences
Abstract. Scientists’ choices of what research topics to pursue are highly consequential and have been the subject of many studies. However, these studies are dispersed across several fields and literatures. This paper provides a review of this body of work. It first reviews theory from economics and sociology to explain how topic choice fits into scientists’ broader competitive strategy, and how the social valuation of research topics reproduces inequalities in science. Then it examines empirical literature on how research topics are chosen in practice, first looking at observational accounts derived from large-scale secondary data sources, then looking at self-reported accounts elicited by surveys and interviews of researchers. Finally, it concludes with a synthesis of the theoretical and empirical literature and identifies research gaps. Themes in these literatures include the rewards associated with certain topics, demographic differences in topic choices, and whether choices are broadly reward maximizing or driven by social context, identities, and path dependence. We also identify several research gaps, including the types and impacts of costs impeding topic entry and switching, the causal mechanisms associating topics with rewards, and the discrepancy between scientists’ self-reported topic motivations and observed behavior.

#predictingthefuture #newfutureofwork | Jaime Teevan
🌱 Prediction: Knowledge will outgrow publication. We’re already seeing academic publication start to buckle under AI, sometimes absurdly. I still publish research more or less the way Darwin did. I run a study, write it up, a few other scientists check it over, and the result gets filed away as a document with my name on the front. Faster than Darwin, with better figures, but the same basic shape. I predict that shape won’t last another decade. Academic authors are starting to slip hidden instructions into papers to flatter the AI that might review them. Reviewers are spending time checking whether citations exist or were hallucinated. Researchers asking AI to tell them about a paper instead of reading it directly. These are signs that the creation of new knowledge is outgrowing the articles that used to contain it. An academic paper serves many purposes at once. It makes an argument legible. It lets strangers check one's reasoning. It assigns credit and responsibility. It records who knew what and when. A paper was the only container we had for these different jobs, so it carried all of them together. With AI, they can be separated. My guess is that means the unit of publication will get smaller. Much of my research has focused on microproductivity, developing the idea that large accomplishments can be built from many small contributions. Publication will start to become a form of microproductivity. Instead of holding onto a result until it can be wrapped in a narrative large enough to justify a paper, researchers will publish it the moment it’s solid. Each finding, method, or negative result will be citable and carry its own provenance, so credit and reasoning travel with it. Reviewing will shrink to match, so claims get checked as they’re made instead of in one verdict at the end. But more than changing publication, the deeper change will be to how research itself is done. You may have heard the term “compound engineering,” where every bug fixed, evaluation written, workflow documented, or lesson learned becomes part of the system’s memory. I predict we’re about to see “compound science,” where every experiment, evaluation, insight, artifact, and learned capability becomes a reusable asset for future discovery. Findings will become evidence. Methods will become building blocks. Failed approaches will become constraints. For centuries, science has relied on humans to navigate an ever-growing body of knowledge. Soon that body of knowledge will help navigate itself. Scientists will spend less time searching for hypotheses and more time deciding which opportunities to pursue. AI systems will propose explanations, design experiments, run analyses, and explore many possibilities in parallel. Every discovery will become a part of the machinery that produces the next one. Papers ten years from now will look less like my current papers than my current papers look like Darwin’s. If they exist at all. #PredictingTheFuture #NewFutureOfWork
Save More Today or Tomorrow: The Role of Urgency in Precommitment Design
To encourage farsighted behaviors, previous research suggests that marketers should invite consumers to precommit to adopting these behaviors “later.” However, the authors propose that people will draw different inferences from different types of precommitment offers, and that these inferences can help explain when precommitment is (and is not) effective at increasing adoption of farsighted behaviors. Specifically, the authors theorize that simultaneously offering consumers the opportunity to adopt a farsighted behavior now or later (i.e., offering “simultaneous precommitment”) may signal that the behavior is not urgently recommended; however, offering consumers the opportunity to adopt that behavior immediately and then, only if they decline, inviting them to adopt it later (i.e., offering “sequential precommitment”) may signal just the opposite. In a multisite field experiment (N = 5,196), the authors find that simultaneously giving consumers the chance to increase their savings now or later reduced retirement savings. Two preregistered lab studies (N = 5,080) show that simultaneous precommitment leads people to infer that taking action is not urgently recommended, and such inferences predict less adoption of recommended behaviors. Importantly, offering sequential precommitment increases inferred urgency, predicting greater adoption. Together, this research advances knowledge about the limits and potential of precommitment.

Project APE: Can policy evaluation be automated? Or is hallucinated slop unavoidable? Let's find out.
An open experiment: AI writes economics papers end-to-end, then competes against peer-reviewed research. Everything public—papers, code, data, failures.
The unintended consequences of large language models as a labor-augmenting technology in science
As a labor-augmenting technology, large language models (LLMs) have the potential to accelerate scientific activity across the research pipeline. But even if LLMs perform on par with human experts at selected tasks, their use will bring unintended consequences as they alter the balance of frictions and inducements that steer the allocation of research effort across projects. Here we develop a simple mathematical model to illustrate. In fields where LLMs are useful primarily as tools for discovering promising projects, researchers will become more selective about what they publish; where they facilitate the process of publishing existing data, researchers will become less selective. By allowing scientists to work more quickly, LLMs raise the opportunity cost of researcher time, creating incentives to refine papers less thoroughly before moving on. Enticing as it is to imagine that, by saving us time on mundane tasks, LLMs will provide us with more time to think deeply and develop projects completely, our results temper such hopes.

The unintended consequences of large language models as a labor-augmenting technology in science
As a labor-augmenting technology, large language models (LLMs) have the potential to accelerate scientific activity across the research pipeline. But even if LLMs perform on par with human experts at selected tasks, their use will bring unintended consequences as they alter the balance of frictions and inducements that steer the allocation of research effort across projects. Here we develop a simple mathematical model to illustrate. In fields where LLMs are useful primarily as tools for discovering promising projects, researchers will become more selective about what they publish; where they facilitate the process of publishing existing data, researchers will become less selective. By allowing scientists to work more quickly, LLMs raise the opportunity cost of researcher time, creating incentives to refine papers less thoroughly before moving on. Enticing as it is to imagine that, by saving us time on mundane tasks, LLMs will provide us with more time to think deeply and develop projects completely, our results temper such hopes.

A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive
Large Language Models (LLMs) are increasingly utilized in autonomous decision-making, where they sample options from vast action spaces. However, the heuristics that guide this sampling process remain under-explored. We study this sampling behavior and show that this underlying heuristics resembles that of human decision-making: comprising a descriptive component (reflecting statistical norm) and a prescriptive component (implicit ideal encoded in the LLM) of a concept. We show that this deviation of a sample from the statistical norm towards a prescriptive component consistently appears in concepts across diverse real-world domains like public health, and economic trends. To further illustrate the theory, we demonstrate that concept prototypes in LLMs are affected by prescriptive norms, similar to the concept of normality in humans. Through case studies and comparison with human studies, we illustrate that in real-world applications, the shift of samples toward an ideal value in LLMs' outputs can result in significantly biased decision-making, raising ethical concerns.
The Last Human-Written Paper: Agent-Native Research Artifacts
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities.

Misperception of Exponential Growth: Are People Aware of Their Errors?
Previous research shows that individuals make systematic errors when judging exponential growth, which has harmful effects for their financial well-being. This study analyzes how far individuals are aware of their errors and how these errors are shaped by arithmetic and conceptual problems. Whereas arithmetic problems could be overcome using computational assistance like a pocket calculator, this is not the case for conceptual problems, a term we use to subsume other error drivers like a general misunderstanding of exponential growth or overwhelming task complexity. In an incentivized experiment, we find that participants strongly overestimate the accuracy of their intuitive judgment. At the same time, their willingness to pay for arithmetic assistance is too high on average, often much above the actual benefits a calculator provides. Using a multitier system of task complexity we can show that the willingness to pay for arithmetic assistance is hardly related to its benefits, indicating that participants do not really understand how the interplay of arithmetic and conceptual problems shape their errors in exponential growth tasks. Our findings are relevant for policymaking and financial advisory practice and can help to design effective approaches to mitigate the detrimental effects of misperceived exponential growth.
