







Abstract. Scientists’ choices of what research topics to pursue are highly consequential and have been the subject of many studies. However, these studies are dispersed across several fields and literatures. This paper provides a review of this body of work. It first reviews theory from economics and sociology to explain how topic choice fits into scientists’ broader competitive strategy, and how the social valuation of research topics reproduces inequalities in science. Then it examines empirical literature on how research topics are chosen in practice, first looking at observational accounts derived from large-scale secondary data sources, then looking at self-reported accounts elicited by surveys and interviews of researchers. Finally, it concludes with a synthesis of the theoretical and empirical literature and identifies research gaps. Themes in these literatures include the rewards associated with certain topics, demographic differences in topic choices, and whether choices are broadly reward maximizing or driven by social context, identities, and path dependence. We also identify several research gaps, including the types and impacts of costs impeding topic entry and switching, the causal mechanisms associating topics with rewards, and the discrepancy between scientists’ self-reported topic motivations and observed behavior.
Why Most Published Research Findings Are False
Summary There is increasing concern that most current published research findings are false. The probability that a research claim is true may depend on study power and bias, the number of other studies on the same question, and, importantly, the ratio of true to no relationships among the relationships probed in each scientific field. In this framework, a research finding is less likely to be true when the studies conducted in a field are smaller; when effect sizes are smaller; when there is a greater number and lesser preselection of tested relationships; where there is greater flexibility in designs, definitions, outcomes, and analytical modes; when there is greater financial and other interest and prejudice; and when more teams are involved in a scientific field in chase of statistical significance. Simulations show that for most study designs and settings, it is more likely for a research claim to be false than true. Moreover, for many current scientific fields, claimed research findings may often be simply accurate measures of the prevailing bias. In this essay, I discuss the implications of these problems for the conduct and interpretation of research.
Shifting the Level of Selection in Science
Criteria for recognizing and rewarding scientists primarily focus on individual contributions. This creates a conflict between what is best for scientists’ careers and what is best for science. In this article, we show how the theory of multilevel selection provides conceptual tools for modifying incentives to better align individual and collective interests. A core principle is the need to account for indirect effects by shifting the level at which selection operates from individuals to the groups in which individuals are embedded. This principle is used in several fields to improve collective outcomes, including animal husbandry, team sports, and professional organizations. Shifting the level of selection has the potential to ameliorate several problems in contemporary science, including accounting for scientists’ diverse contributions to knowledge generation, reducing individual-level competition, and promoting specialization and team science. We discuss the difficulties associated with shifting the level of selection and outline directions for future development in this domain.

The Discovery Engine: A Framework for AI-Driven Synthesis and Navigation of Scientific Knowledge Landscapes
Scientific progress relies on the effective accumulation, synthesis, and critical evaluation of knowledge. Traditionally, the well-documented, peer reviewed publication served as the primary standard for filtering and disseminating credible findings within the scientific community. Recently, however, we are witnessing an unprecedented acceleration in research output, a veritable explosion of scientific publications across all disciplines [1]. Yet, this very abundance creates a paradox: the sheer volume threatens to overwhelm the mechanisms designed for its assimilation and synthesis. Researchers, even within highly specialized subfields, face an almost insurmountable challenge in keeping abreast of relevant developments, integrating disparate findings, and identifying the truly novel signals amidst the noise [2]. This information overload contributes to disciplinary fragmentation, hindering the cross-pollination of ideas essential for disruptive innovation [3]. Furthermore, persistent concerns regarding "reproducibility crisis" [2], predatory journals, inflation of research areas[4], growing retractions and the potential influences of bibliometrics on research direction [5] highlight systemic challenges in validating and prioritizing scientific contributions to fundamental knowledge.
Why we built the Journal of Research on Research and what it cost us - LSE Impact
Gemma Derrick, Bart Penders, Serge Horbach & Tony Ross-Hellauer reflect on the choices, challenges & compromises of launching the Journal of Research on Research

A billion-dollar donation: estimating the cost of researchers’ time spent on peer review
The amount and value of researchers’ peer review work is critical for academia and journal publishing. However, this labor is under-recognized, its magnitude is unknown, and alternative ways of organizing peer review labor are rarely considered.

Do Science <i>Kardashians</i> Get Citation Premium? Self‐Fulfilling Effects of Social Media on Scientific Impact
ABSTRACT We analyze whether the visibility of scientists on social media affects the number of academic citations. We use the global COVID‐19 pandemic as a quasinatural experiment that exogenously increased public attention and the demand for expertise. Using publications on COVID‐related topics by social media stars and their coauthors prior to the outbreak of the pandemic, we find that social media stars' pre‐COVID‐era papers received about – more citations annually per paper after 2019. Quantitatively comparable results are obtained when we use scientists' Kardashian index (K‐index) as a benchmark for stardom, however we find no significant effects when using the intensive margin of scientists' K‐indexes. We provide a brief discussion of policy implications in light of these findings.

The unintended consequences of large language models as a labor-augmenting technology in science
As a labor-augmenting technology, large language models (LLMs) have the potential to accelerate scientific activity across the research pipeline. But even if LLMs perform on par with human experts at selected tasks, their use will bring unintended consequences as they alter the balance of frictions and inducements that steer the allocation of research effort across projects. Here we develop a simple mathematical model to illustrate. In fields where LLMs are useful primarily as tools for discovering promising projects, researchers will become more selective about what they publish; where they facilitate the process of publishing existing data, researchers will become less selective. By allowing scientists to work more quickly, LLMs raise the opportunity cost of researcher time, creating incentives to refine papers less thoroughly before moving on. Enticing as it is to imagine that, by saving us time on mundane tasks, LLMs will provide us with more time to think deeply and develop projects completely, our results temper such hopes.

The unintended consequences of large language models as a labor-augmenting technology in science
As a labor-augmenting technology, large language models (LLMs) have the potential to accelerate scientific activity across the research pipeline. But even if LLMs perform on par with human experts at selected tasks, their use will bring unintended consequences as they alter the balance of frictions and inducements that steer the allocation of research effort across projects. Here we develop a simple mathematical model to illustrate. In fields where LLMs are useful primarily as tools for discovering promising projects, researchers will become more selective about what they publish; where they facilitate the process of publishing existing data, researchers will become less selective. By allowing scientists to work more quickly, LLMs raise the opportunity cost of researcher time, creating incentives to refine papers less thoroughly before moving on. Enticing as it is to imagine that, by saving us time on mundane tasks, LLMs will provide us with more time to think deeply and develop projects completely, our results temper such hopes.

The strain on scientific publishing
Scientists are increasingly overwhelmed by the volume of articles being published. Total articles indexed in Scopus and Web of Science have grown exponentially in recent years; in 2022 the article total was approximately ~47% higher than in 2016, which has outpaced the limited growth - if any - in the number of practising scientists. Thus, publication workload per scientist (writing, reviewing, editing) has increased dramatically. We define this problem as the strain on scientific publishing. To analyse this strain, we present five data-driven metrics showing publisher growth, processing times, and citation behaviours. We draw these data from web scrapes, requests for data from publishers, and material that is freely available through publisher websites. Our findings are based on millions of papers produced by leading academic publishers. We find specific groups have disproportionately grown in their articles published per year, contributing to this strain. Some publishers enabled this growth by adopting a strategy of hosting special issues, which publish articles with reduced turnaround times. Given pressures on researchers to publish or perish to be competitive for funding applications, this strain was likely amplified by these offers to publish more articles. We also observed widespread year-over-year inflation of journal impact factors coinciding with this strain, which risks confusing quality signals. Such exponential growth cannot be sustained. The metrics we define here should enable this evolving conversation to reach actionable solutions to address the strain on scientific publishing.

How Field Experiments in Economics Can Complement Psychological Research on Judgment Biases
This review summarizes results of field experiments examining individual behaviors across several market settings—from open-air markets to rideshare markets to tax-compliance markets—where people sort themselves into market roles wherein they make consequential decisions. Using three distinct examples from my own research on the endowment effect, left-digit bias, and omission bias, I showcase how field experiments can help researchers understand mediators, heterogeneity, and causal moderation involved in judgment biases in the field. In this manner, the review highlights that economic field experiments can serve an invaluable intellectual role alongside traditional laboratory research.

In Praise of Moderation: Suggestions for the Scope and Use of Pre-Analysis Plans for RCTs in Economics
Pre-Analysis Plans (PAPs) for randomized evaluations are becoming increasingly common in Economics, but their definition remains unclear and their practical applications therefore vary widely. Based on our collective experiences as researchers and editors, we articulate a set of principles for the ex-ante scope and ex-post use of PAPs. We argue that the key benefits of a PAP can usually be realized by completing the registration fields in the AEA RCT Registry. Specific cases where more detail may be warranted include when subgroup analysis is expected to be particularly important, or a party to the study has a vested interest. However, a strong norm for more detailed pre-specification can be detrimental to knowledge creation when implementing field experiments in the real world. An ex-post requirement of strict adherence to pre-specified plans, or the discounting of non-pre-specified work, may mean that some experiments do not take place, or that interesting observations and new theories are not explored and reported. Rather, we recommend that the final research paper be written and judged as a distinct object from the “results of the PAP”; to emphasize this distinction, researchers could consider producing a short, publicly available report (the “populated PAP”) that populates the PAP to the extent possible and briefly discusses any barriers to doing so.

Faster science, penalties in evaluation, and concerns on quality and impact: Researchers’ use and perceptions of preprints
The preprint ecosystem has expanded rapidly over the past decade, fundamentally altering science communication. Yet, the scholarly community’s attitudes toward this shift remain underexplored. Through a large-scale survey of US and Canadian biomedical scholars, we provide a comprehensive analysis of preprint utilization, perceived impact, and integration into academic credit systems. We find robust engagement across reading, citing, and submitting preprints; however, this activity is driven primarily by a desire for rapid dissemination rather than a foundational commitment to open science. Furthermore, while preprints are valued as networking assets, perceived career penalties during formal academic evaluations stifle broader cultural adoption. Crucially, to navigate the absence of formal peer review, scholars report a heavy reliance on author reputation as a primary heuristic to evaluate a preprint’s credibility and guide their reading and citation decisions. Notably, despite acknowledging preprints’ role in accelerating knowledge sharing, scholars express significant concerns regarding fraud and misinformation, particularly amid declining public trust in science and emerging threats to scientific integrity from artificial intelligence. To resolve these tensions, the preprint ecosystem must evolve beyond prioritizing speed to foster genuine academic dialogue. Simultaneously, evaluation frameworks must adapt to the realities of preprinting, and innovative quality-control mechanisms are urgently needed to balance rapid dissemination with rigorous scientific integrity.

Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection
Large language models (LLMs) are rapidly being adopted as research assistants, particularly for literature review and reference recommendation, yet little is known about whether they introduce demographic bias into citation workflows. This study systematically investigates gender bias in LLM-driven reference selection using controlled experiments with pseudonymous author names. We evaluate several LLMs (GPT-4o, GPT-4o-mini, Claude Sonnet, and Claude Haiku) by varying gender composition within candidate reference pools and analyzing selection patterns across fields. Our results reveal two forms of bias: a persistent preference for male-authored references and a majority-group bias that favors whichever gender is more prevalent in the candidate pool. These biases are amplified in larger candidate pools and only modestly attenuated by prompt-based mitigation strategies. Field-level analysis indicates that bias magnitude varies across scientific domains, with social sciences showing the least bias. Our findings indicate that LLMs can reinforce or exacerbate existing gender imbalances in scholarly recognition. Effective mitigation strategies are needed to avoid perpetuating existing gender disparities in scientific citation practices before integrating LLMs into high-stakes academic workflows.

Screening, sorting, and the feedback cycles that imperil peer review
Scholarly journals rely on peer review to identify the science most worthy of publication. Yet finding willing and qualified reviewers to evaluate manuscripts has become an increasingly challenging task, possibly even threatening the long-term viability of peer review as an institution. What can or should be done to salvage it? Here, we develop mathematical models to reveal the intricate interactions among incentives faced by authors, reviewers, and readers in their endeavors to identify the best science. Two facets are particularly salient. First, peer review partially reveals authors’ private sense of their work’s quality through their decisions of where to send their manuscripts. Second, journals’ reliance on traditionally unpaid and largely unrewarded review labor deprives them of a standard market mechanism—wages—to recruit additional reviewers when review labor is in short supply. We highlight a resulting feedback loop that threatens to overwhelm the peer review system: (1) an increase in submissions overtaxes the pool of suitable peer reviewers; (2) the accuracy of review drops because journals must either solicit assistance from less qualified reviewers or ask current reviewers to do more; (3) as review accuracy drops, submissions further increase as more authors try their luck at venues that might otherwise be a stretch. We illustrate how this cycle is propelled by the increasing emphasis on high-impact publications, the proliferation of journals, and competition among these journals for peer reviews. Finally, we suggest interventions that could slow or even reverse this cycle of peer-review meltdown.
Characterizing the causes, dynamics, and consequences of choice deferral
The fact that people often avoid making decisions is well known, and past research has helped to identify some of the conditions and reasons for doing so. For instance, people may forgo choices between bad options because they prefer not to end up with one of those options. It is much less clear when and why people avoid choosing in cases where they eventually will have to make a given decision. To study such instances of choice deferral, we presented participants with a series of choices, and for each choice they were allowed to either choose immediately or defer the decision until later in the experiment. Across six experiments and three choice domains (choices among consumer goods, artwork, and political candidates), we find that the strongest predictor of choice deferral is the overall value of a given set of options, with relative value (i.e., how hard it is to identify the best option) counterintuitively playing a smaller role. We show that the influence of overall value on choice deferral can be accounted for by a dynamic decision model according to which participants appraise the option set relative to a criterion before deciding whether to choose or defer, comparing this to a previous model whereby participants make such a decision based on a predetermined decision time limit. We further reveal that the influence of overall value on choice deferral is determined by how congruent options are with a given choice goal (choose-best or choose-worst) rather than simply how bad those options are. Collectively, our findings shed new light on how people decide to put off the inevitable.
Against theory-motivated experimentation: Can random experimental choice lead to better theories?
Scientists must choose which among many experiments to perform. We study the epistemic success of experimental choice strategies proposed by philosophers of science or executed by scientists themselves. We develop a multi-agent model of the scientific process that jointly formalizes its core aspects: active experimentation, theorizing, and social learning. We find that agents who choose new experiments at random develop the most informative and predictive theories of the world. The agents aiming to confirm, falsify theories, or resolve theoretical disagreements end up with an illusion of epistemic success: they develop promising accounts for the data they collected, while misrepresenting the ground truth that they intended to learn about. Agents experimenting in these theory-motivated ways acquire less diverse or less representative samples from the ground truth that also turn out to be easier to account for. Random data collection, on the other hand, combines virtues of diverse and representative sampling from a target scientific domain which enables cumulative development of the successful theoretical accounts of it. We suggest that randomization, already a gold standard within experiments, is also beneficial at the level of experiments themselves.
