







Scientists must choose which among many experiments to perform. We study the epistemic success of experimental choice strategies proposed by philosophers of science or executed by scientists themselves. We develop a multi-agent model of the scientific process that jointly formalizes its core aspects: active experimentation, theorizing, and social learning. We find that agents who choose new experiments at random develop the most informative and predictive theories of the world. The agents aiming to confirm, falsify theories, or resolve theoretical disagreements end up with an illusion of epistemic success: they develop promising accounts for the data they collected, while misrepresenting the ground truth that they intended to learn about. Agents experimenting in these theory-motivated ways acquire less diverse or less representative samples from the ground truth that also turn out to be easier to account for. Random data collection, on the other hand, combines virtues of diverse and representative sampling from a target scientific domain which enables cumulative development of the successful theoretical accounts of it. We suggest that randomization, already a gold standard within experiments, is also beneficial at the level of experiments themselves.
Mechanism Experiments and Policy Evaluations
Randomized controlled trials are increasingly used to evaluate policies. How can we make these experiments as useful as possible for policy purposes? We argue greater use should be made of experiments that identify the behavioral mechanisms that are central to clearly specified policy questions, what we call "mechanism experiments." These types of experiments can be of great policy value even if the intervention that is tested (or its setting) does not correspond exactly to any realistic policy option.
Eliciting Beliefs with Random Generation Tasks
Elicitation methods, such as asking people to produce the deciles of a distribution, are standard practices in policy or applied statistics. Similarly, much of cognitive science and psychology focuses on determining people's people's beliefs or latent traits through questionnaires or judgment tasks. However, these approaches often only capture a rough outline of what people know and are usually limited to point estimates of people's beliefs. Here, we present a novel experimental paradigm that allows us to access people's beliefs and how variable these beliefs are. Our task is based on an established random generation paradigm in which participants produce quantities from a particular domain as randomly as possible. We hypothesize that due to the minds' general-purpose mechanisms for probabilistic inferences, these random sequences represent the participants' underlying prior beliefs. We show that our method can infer participants' beliefs for a wide range of numeric quantities at comparable accuracy as an established elicitation method. Moreover, these inferred beliefs are consistent with individual participants' generalization and inference patterns in a subsequent conditional prediction task. We then extend our approach to non-numeric belief elicitation, highlighting how our method can go beyond numeric elicitation and provide insight into complex beliefs that are challenging to assess experimentally. Empirically, our results highlight that people know the rough shapes of environmental distributions, and these beliefs guide inference and generalization. Moreover, using our novel approach, we also show that people know the fine details of environmental distributions. Finally, our experimental results show that random generation paradigms can be a useful tool for cognitive scientists, psychologists, and applied statisticians.
Underspecified Human Decision Experiments Considered Harmful
Decision-making with information displays is a key focus of research in areas like human-AI collaboration and data visualization. However, what constitutes a decision problem, and what is required for an experiment to conclude that decisions are flawed, remain imprecise. We present a widely applicable definition of a decision problem synthesized from statistical decision theory and information economics. We claim that to attribute loss in human performance to bias, an experiment must provide the information that a rational agent would need to identify the normative decision. We evaluate whether recent empirical research on AI-assisted decisions achieves this standard. We find that only 10 (26%) of 39 studies that claim to identify biased behavior presented participants with sufficient information to make this claim in at least one treatment condition. We motivate the value of studying well-defined decision problems by describing a characterization of performance losses they allow to be conceived.

Philosophy Experiments — Interactive Thought Experiments
Explore 27 interactive philosophy experiments that challenge your moral intuitions, test your logical reasoning, and probe the boundaries of belief.

Artificial intelligence and illusions of understanding in scientific research
Scientists are enthusiastically imagining ways in which artificial intelligence (AI) tools might improve research. Why are AI tools so attractive and what are the risks of implementing them across the research pipeline? Here we develop a taxonomy of scientists’ visions for AI, observing that their appeal comes from promises to improve productivity and objectivity by overcoming human shortcomings. But proposed AI solutions can also exploit our cognitive limitations, making us vulnerable to illusions of understanding in which we believe we understand more about the world than we actually do. Such illusions obscure the scientific community’s ability to see the formation of scientific monocultures, in which some types of methods, questions and viewpoints come to dominate alternative approaches, making science less innovative and more vulnerable to errors. The proliferation of AI tools in science risks introducing a phase of scientific enquiry in which we produce more but understand less. By analysing the appeal of these tools, we provide a framework for advancing discussions of responsible knowledge production in the age of AI.

The High Cost of Not Doing Experiments - Behavioral Scientist
We can pay dearly, in blood, treasure, and well-being, for experiments that aren’t done. -Richard Nisbett, Mindware

Illusions of Understanding in the Sciences
Scientists seek to understand the causes of observed phenomena. Beliefs that they have succeeded are based on understanding that is rarely or possibly never complete, and varies in depth and quality. Most often scientists believe they understand more than they do, making their belief an illusion. This illusion then persists in explanations scientists provide in print, in talks, or in discussions. The illusion that a scientist has a valid and complete explanation tends to be magnified when the data are well described by mathematical and computer simulation models due to the precision of such models and their ability to predict well; prediction does not imply causality, but gives the illusion that it does. The first part of this essay supports the case for the universality of partial and incomplete levels of understanding by showing the difficulty of reaching a deep level of understanding for even a simple analysis and model that most scientists use and believe they understand: linear regression. The second part highlights some implications of the existence of many levels of understanding and explanation, and their use by scientists for design, testing, analysis, and theory development. It discusses the way that deduction and induction depend on the levels of understanding and the implications of the illusion that a scientist’s understanding is deep. It makes a case that the many incomplete levels of understanding affect, often unwittingly, the ways scientists design experiments, test theories, comprehend, communicate, and teach.

Social learning preserves both useful and useless theories by canalizing learners’ exploration
In many domains, learning from others is crucial for leveraging cumulative cultural knowledge, which encapsulates the efforts of successive generations of innovators. However, anecdotal and experimental evidence suggests that reliance on social information can reduce the exploration of the problem space. Here, we experimentally investigate the extent to which cultural transmission fosters the persistence of arbitrary solutions in a context where participants are incentivized to improve a physical system across multiple trials. Participants were exposed to various theories about the system, ranging from accurate to misleading. Our findings indicate that even under conditions conducive to exploration, the transmission of cultural knowledge canalizes learners’ focus, limiting their consideration of alternative solutions. This effect was observed in both the theories produced and the solutions attempted by participants, irrespective of the accuracy of the provided theories. These results challenge the notion that arbitrary solutions persist only when they are efficient or intuitive and underscore the significant role of cultural transmission in shaping human knowledge and technologies.

Can AI Agents Synthesize Scientific Conclusions?
Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes domains such as health remains unclear. We introduce SciConBench, a large-scale live benchmark of 9.11K questions and expert-written conclusions from systematic reviews to evaluate open-domain scientific conclusion synthesis. The benchmark draws on an expert-validated automated evaluation pipeline that decomposes conclusions into atomic facts and measures correctness and comprehensiveness via factual precision and recall. To mitigate data leakage, we further introduce SciConHarness, a clean-room evaluation harness that equips agents with controlled web interaction to ensure valid measurement. Evaluating 8 frontier models and deep research agents, we find that factual quality remains low: under clean-room settings, the best agent achieves only a factual F1 of 0.337. Our clean-room setting consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities. Finally, we audit consumer-facing agents (e.g., Google AI Overview, OpenEvidence) and find they frequently generate incomplete and sometimes contradictory conclusions, even when the ground-truth answer is available. Overall, our results show that reliable synthesis of scientific conclusions remains an open challenge, and that clean-room evaluation is essential for assessing open-domain AI agents.

In Praise of Moderation: Suggestions for the Scope and Use of Pre-Analysis Plans for RCTs in Economics
Pre-Analysis Plans (PAPs) for randomized evaluations are becoming increasingly common in Economics, but their definition remains unclear and their practical applications therefore vary widely. Based on our collective experiences as researchers and editors, we articulate a set of principles for the ex-ante scope and ex-post use of PAPs. We argue that the key benefits of a PAP can usually be realized by completing the registration fields in the AEA RCT Registry. Specific cases where more detail may be warranted include when subgroup analysis is expected to be particularly important, or a party to the study has a vested interest. However, a strong norm for more detailed pre-specification can be detrimental to knowledge creation when implementing field experiments in the real world. An ex-post requirement of strict adherence to pre-specified plans, or the discounting of non-pre-specified work, may mean that some experiments do not take place, or that interesting observations and new theories are not explored and reported. Rather, we recommend that the final research paper be written and judged as a distinct object from the “results of the PAP”; to emphasize this distinction, researchers could consider producing a short, publicly available report (the “populated PAP”) that populates the PAP to the extent possible and briefly discusses any barriers to doing so.

AI Research Agents Narrow Scientific Exploration
AI research agents can now generate research ideas, design experiments, run code, and draft papers, raising the possibility of large-scale AI-assisted scientific discovery. Many current agent frameworks explicitly encourage the generation of novel and high-impact ideas. Yet it remains unclear whether AI-assisted ideation broadens scientific exploration or mainly concentrates around existing work. We study AI research agents as scientific search systems. Using four AI research-agent frameworks and six large language models, we generate 37,802 scientific ideas from shared seed literature across citation-defined research areas in AI and machine learning. We then compare the resulting AI ideas against human-authored papers from the same research areas, follow-on human research emerging from the same seed literature, and the seed literature itself. Across experiments, four consistent patterns emerge. First, AI-generated ideas are substantially more concentrated than human-authored papers from the same research areas. Second, AI-generated ideas remain much closer to their starting literature than later human follow-on work does. Third, papers most similar to AI-generated ideas tend to receive lower subsequent citations. Fourth, when AI-generated ideas differ from prior work, the differences arise primarily from recombining existing technical methods rather than introducing fundamentally new research questions. Overall, current AI research agents appear better suited to local elaboration than to broadening scientific exploration.

From Predictive Algorithms to Automatic Generation of Anomalies
Machine learning algorithms can find predictive signals that researchers fail to notice; yet they are notoriously hard-to-interpret. How can we extract theoretical insights from these black boxes? History provides a clue. Facing a similar problem – how to extract theoretical insights from their intuitions – researchers often turned to “anomalies:” constructed examples that highlight flaws in an existing theory and spur the development of new ones. Canonical examples include the Allais paradox and the Kahneman-Tversky choice experiments for expected utility theory. We suggest anomalies can extract theoretical insights from black box predictive algorithms. We develop procedures to automatically generate anomalies for an existing theory when given a predictive algorithm. We cast anomaly generation as an adversarial game between a theory and a falsifier, the solutions to which are anomalies: instances where the black box algorithm predicts - were we to collect data - we would likely observe violations of the theory. As an illustration, we generate anomalies for expected utility theory using a large, publicly available dataset on real lottery choices. Based on an estimated neural network that predicts lottery choices, our procedures recover known anomalies and discover new ones for expected utility theory. In incentivized experiments, subjects violate expected utility theory on these algorithmically generated anomalies; moreover, the violation rates are similar to observed rates for the Allais paradox and Common ratio effect.

Re-Engineering Wimsatt for Limited Beings
Science is the best way to produce facts about reality. The best, at least, that limited human beings have devised so far. Yet, not even scientists quite seem to understand how scientific knowledge is generated. This is not only a philosophical but also a practical problem, as our misunderstandings affect the quality of our research and limit the directions it can take. In light of this, it may be good if we reflected a bit more on how we do science — to become better researchers through philosophy. Here, I provide an accessible introduction to a philosophical approach that achieves precisely this: William Wimsatt’s multi-perspectival realism. It disabuses us of widespread but misleading myths and idealizations about science, such as the idea that everything in the world can be reduced to a fundamental level, or that we can approach a “view from nowhere” — complete and objectively detached knowledge of the world. Wismatt proposes an alternative view based on his thorough studies of actual research practice. It cuts deeply into the layered yet messy structure of reality, and the improvised but potent tools we have available, as limited and evolved beings, to explore it. Wimsatt reframes science as an irregular yet adaptive process rather than a cumulative repository of unalterable facts. His philosophy provides a workable and grounded middle way between radical skepticism and naïve belief in the objective truth of science. It explains how knowledge is conceptually constructed by humans, but still connects us to reality in a trustworthy way. We need such a new view of science, not only to improve our research practices and outcomes but, more generally, to gain a more realistic understanding of ourselves, the world, and our place and role within it.

The Law of Conservation of Information: Search Processes Only Redistribute Existing Information
Conservation of information sparked scientific interest once a recurring pattern was noticed in the evolutionary computing literature. In grappling with the creation of information through evolutionary algorithms, this literature consistently revealed that the information outputted by such algorithms always needed first to be programmed into them. Thus, the primary goal of this literature—to uncover how information could be created from scratch or de novo —was shown to be misconceived: the information was not created but instead shuffled around or smuggled in, implying that it already existed in some form or other. Information output in these situations therefore always presupposed a counterbalancing input of prior information. Once this pattern was seen, the next logical step was to quantify the amount of information inputted and outputted, demonstrating a consistent mathematical relation between the two. This led to the proof of a number of theorems about search. In these theorems, a baseline search with probability p of success gave way to an improved search with probability q of success. Typically p would be very small and close to zero, implying a practically impossible search (like searching for a needle in a haystack). By contrast, q would be much larger and close to one, implying an eminently doable search. The punchline of these theorems was that, as the improved search became itself the subject of a search (a search for a search , or S4S), the probability of finding it could not exceed p / q , rendering success of the improved search no more probable than success of the original baseline search, in effect filling one hole by digging another. Such conservation-of-information theorems, as they came to be called, were search-space specific, adapted to different kinds of search across a range of search spaces. There was a measure-theoretic theorem in which probability measures guided search. There were also function-theoretic and fitness-theoretic theorems where mappings into the search space as well as fitness functions on the search space respectively guided search. The key insight of this paper is that all these conservation-of-information theorems are special cases of a simple probabilistic relation based on elementary probability theory. This paper identifies the underlying rationale that makes all the previous conservation-of-information theorems work. In so doing, it provides a straightforward proof and general formulation of what may rightly be called the Law of Conservation of Information.
Why Most Published Research Findings Are False
Summary There is increasing concern that most current published research findings are false. The probability that a research claim is true may depend on study power and bias, the number of other studies on the same question, and, importantly, the ratio of true to no relationships among the relationships probed in each scientific field. In this framework, a research finding is less likely to be true when the studies conducted in a field are smaller; when effect sizes are smaller; when there is a greater number and lesser preselection of tested relationships; where there is greater flexibility in designs, definitions, outcomes, and analytical modes; when there is greater financial and other interest and prejudice; and when more teams are involved in a scientific field in chase of statistical significance. Simulations show that for most study designs and settings, it is more likely for a research claim to be false than true. Moreover, for many current scientific fields, claimed research findings may often be simply accurate measures of the prevailing bias. In this essay, I discuss the implications of these problems for the conduct and interpretation of research.
Integrated information and predictive processing theories of consciousness: An adversarial collaborative review
As neuroscientific theories of consciousness continue to proliferate, the need to assess their similarities and differences - as well as their predictive and explanatory power - becomes ever more pressing. Recently, a number of structured adversarial collaborations have been devised to test the competing predictions of several candidate theories of consciousness. In this review, we compare and contrast three theories being investigated in one such adversarial collaboration: Integrated Information Theory, Neurorepresentationalism, and Active Inference. We begin by presenting the core claims of each theory, before comparing them in terms of the phenomena they seek to explain, the sorts of explanations they avail, and the methodological strategies they endorse. We then consider some of the inherent challenges of theory-testing, and how adversarial collaboration addresses some of these difficulties. The stage is then set for the empirical work to come: first, we outline the key hypotheses to be tested across a series of multi-site experiments; second, we discuss the kinds of observations that would support or challenge each theory; third, we consider how these theories might assimilate or accommodate such observations. Finally, we show how data harvested across disparate experiments (and their replicates) may be formally integrated to provide a quantitative measure of the evidential support accrued under each theory. Besides orienting the reader to the theoretical foundations of our collaboration, this review aims to provide valuable meta-scientific insights into the mechanics of adversarial collaboration and theory-testing in general - including the way theories may be evaluated in terms of the scientific progress they deliver.
