







A recurrent theme in computational science (and elsewhere) is the need to combine machine-readable information (which in the following I will call "facts" for simplicity) with a narrative for the benefit of human readers. The most obvious situation is a scientific publication, which is essentially a narrative explaining the context and motivation for a study, the work that was undertaken, the results that were observed, and conclusions drawn from these results. For a scientific study that made use of computation (which is almost all of today's research work), the narrative refers to various computational facts, in particular machine-readable input data, program code, and computed results.
Zechen Zhang on Twitter / X
1/ For nearly 350 years, science has communicated itself through one object: the paper. A linear narrative, frozen as a PDF, written for a human reader. We've come to treat that format as the medium of science itself.It doesn't have to be. It's a historical artifact. 🧵 pic.twitter.com/P5UUhceLkC— Zechen Zhang (@ZechenZhang5) April 30, 2026

The Last Human-Written Paper: Agent-Native Research Artifacts
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities.

Konrad Hinsen's blog
Thirty years after my first contact with computational (ir)reproducibility, I am happy to note that many things have improved. Reproducibility, computational and otherwise, is increasingly recognized as an important aspect of scientific quality control, and mostly considered worth striving for. However, I also note that more and more people, including reproducibility activists, have lost contact with the day-to-day reality in which reproducibility matters. Reproducibility is becoming an item on a checklist, and its precise incarnation the subject of political bickering aimed at making it easy to check off that item. So let's take a look at why computational reproducibility matters for researchers.
The new way we’ll do science
Papers should become human-readable views over a graph of data, tools, results, and certificates.

Konrad Hinsen's blog
Knowledge refinement is the ever ongoing process in science (and beyond it) that shepherds knowledge from lab notebooks into journal articles and then on to review articles, monographs, reference handbooks, university textbooks, and finally professional domain expertise and school education for a wider public. It has been going on for a few centuries, but we hardly talk about it. In fact, I made up the term because I couldn't find an established one. Computational knowledge has not yet found its place in the knowledge refinement process. Why not? And what can we do to make it happen?
What stories should we tell about scientific progress now?
In search of the mysterious fruits of basic science

#predictingthefuture #newfutureofwork | Jaime Teevan
🌱 Prediction: Knowledge will outgrow publication. We’re already seeing academic publication start to buckle under AI, sometimes absurdly. I still publish research more or less the way Darwin did. I run a study, write it up, a few other scientists check it over, and the result gets filed away as a document with my name on the front. Faster than Darwin, with better figures, but the same basic shape. I predict that shape won’t last another decade. Academic authors are starting to slip hidden instructions into papers to flatter the AI that might review them. Reviewers are spending time checking whether citations exist or were hallucinated. Researchers asking AI to tell them about a paper instead of reading it directly. These are signs that the creation of new knowledge is outgrowing the articles that used to contain it. An academic paper serves many purposes at once. It makes an argument legible. It lets strangers check one's reasoning. It assigns credit and responsibility. It records who knew what and when. A paper was the only container we had for these different jobs, so it carried all of them together. With AI, they can be separated. My guess is that means the unit of publication will get smaller. Much of my research has focused on microproductivity, developing the idea that large accomplishments can be built from many small contributions. Publication will start to become a form of microproductivity. Instead of holding onto a result until it can be wrapped in a narrative large enough to justify a paper, researchers will publish it the moment it’s solid. Each finding, method, or negative result will be citable and carry its own provenance, so credit and reasoning travel with it. Reviewing will shrink to match, so claims get checked as they’re made instead of in one verdict at the end. But more than changing publication, the deeper change will be to how research itself is done. You may have heard the term “compound engineering,” where every bug fixed, evaluation written, workflow documented, or lesson learned becomes part of the system’s memory. I predict we’re about to see “compound science,” where every experiment, evaluation, insight, artifact, and learned capability becomes a reusable asset for future discovery. Findings will become evidence. Methods will become building blocks. Failed approaches will become constraints. For centuries, science has relied on humans to navigate an ever-growing body of knowledge. Soon that body of knowledge will help navigate itself. Scientists will spend less time searching for hypotheses and more time deciding which opportunities to pursue. AI systems will propose explanations, design experiments, run analyses, and explore many possibilities in parallel. Every discovery will become a part of the machinery that produces the next one. Papers ten years from now will look less like my current papers than my current papers look like Darwin’s. If they exist at all. #PredictingTheFuture #NewFutureOfWork
Konrad Hinsen's blog
A much cited essay by Bret Victor, "Explorable Explanations", argues for supporting and encouraging active reading in communicating ideas. Explanatory text should thus be complemented by interactive visualizations and computational demonstrations, allowing the reader to actively engage with the ideas. If you haven't read Victor's essay yet, please do so now, and then come back here. It's not very long. What I am going to discuss is a variation on Victor's proposal, and I won't repeat his well-presented arguments.
Illusions of Understanding in the Sciences
Scientists seek to understand the causes of observed phenomena. Beliefs that they have succeeded are based on understanding that is rarely or possibly never complete, and varies in depth and quality. Most often scientists believe they understand more than they do, making their belief an illusion. This illusion then persists in explanations scientists provide in print, in talks, or in discussions. The illusion that a scientist has a valid and complete explanation tends to be magnified when the data are well described by mathematical and computer simulation models due to the precision of such models and their ability to predict well; prediction does not imply causality, but gives the illusion that it does. The first part of this essay supports the case for the universality of partial and incomplete levels of understanding by showing the difficulty of reaching a deep level of understanding for even a simple analysis and model that most scientists use and believe they understand: linear regression. The second part highlights some implications of the existence of many levels of understanding and explanation, and their use by scientists for design, testing, analysis, and theory development. It discusses the way that deduction and induction depend on the levels of understanding and the implications of the illusion that a scientist’s understanding is deep. It makes a case that the many incomplete levels of understanding affect, often unwittingly, the ways scientists design experiments, test theories, comprehend, communicate, and teach.

LLM use in scholarly writing poses a provenance problem
Nature Machine Intelligence - LLM use in scholarly writing poses a provenance problem

Konrad Hinsen's blog
By now, most scientists have probably seen figures, tables, and even entire journal articles made by so-called "generative AI", containing more or less subtle mistakes or inconsistencies. What I haven't seen yet, but expect to see soon, is the scientific equivalent of deepfakes: made-up results that come with made-up code that reproduces them. This is likely to become a new challenge for reproducible research.
Introducing Claude Fable 5.1 and Claude Mythos 5.1
Our most advanced models for coding and knowledge work. Their research capabilities also offer an early glimpse of how AI models will contribute to scientific progress.


Steps Towards an Infrastructure for Scholarly Synthesis
Sharing, reusing, and synthesizing knowledge is central to the research process, both individually, and with others. These core functions are not supported by our formal scholarly publishing infrastructure: instead of the smooth functioning of functional infrastructure, researchers resort to laborious "hacks" and workarounds to "mine" publications for what they need, and struggle to efficiently share the resulting information with others. Information scientists have proposed an alternative infrastructure based on the more appropriately granular model of a discourse graph of claims, and evidence, along with key rhetorical relationships between them. However, despite significant technical progress on standards and platforms, the predominant infrastructure remains steadfastly document-based. Drawing from infrastructure studies, we locate the current infrastructural bottlenecks in the lack of local systems that integrate discourse-centric models to augment synthesis work, from which an infrastructure for synthesis can be grown. Through 3 years of research through design and field deployment in a distributed community of hypertext notebook users, we elaborate a design vision of what can and should be built in order to grow a discourse-centric synthesis infrastructure: a thriving "installed base" of researchers authoring local, shareable discourse graphs to improve synthesis work, enhance primary research and research training, and augment collaborative research. We discuss how this design vision -- and our empirical work -- contributes steps towards a new infrastructure for synthesis, and increases HCI's capacity to advance collective intelligence and solve infrastructure-level problems.

Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers
The concept of reproducibility can have different interpretations across various research fields and even within the same field [39]. To avoid confusion, we first specify our terms, broadly defining reproducibility and then further categorizing it into various types and degrees. The first distinction comes from Goodman et al. [42], who specify a fundamental division between whether we (i) mean reproducible in principle (termed “methods” reproducibility) due to sufficient description/sharing of methodologies, materials, etc., or (ii) whether results/conclusions actually prove to be reproducible when experiments or analyses are re-done. In the second category, they distinguish “results” and “inferential” reproducibility, depending on whether the analyses or inferences to broader conclusions are reproduced.
Konrad Hinsen's blog
Home Page - Software Heritage
GNU Guix transactional package manager and distribution — GNU Guix

Keynote: Reproducibility and replicability of computer simulations | Canal U
Reproducible research: methodological principles for transparent…
Reproducible Research II: Practices and tools for managing compu…