







At the recent SciCodes Symposium, I brought up the question of reviewing research software during the panel discussion. One panelist then raised the question of why we should review research software. I found this question surprising at first, but I do agree that it deserves an answer. Here is mine.
Science is open software
TL;DR I claim that modern science is synonymous with open source software. This post explains why, why it matters, and what you can (and should) do next. Why do you care about (open source) software? - Everyone I spend a lot of my time working on software. I have been asked why software matters more times than I can remember. Software is, people say, not science. It’s a time sink, something to rush past in the pursuit of what really matters: results (and papers if you’re in academia).
Linking the world's research to the code it runs on - OpenAlex blog
Research relies on software. Software written by scientists, for science, runs through the entire modern research stack: NumPy and SciPy, R and ggplot2, Jupyter, BLAST, ImageJ, AlphaFold. Yet in the scholarly record, that software is nearly invisible. Software is not usually cited formally in publications and is usually just mentioned in the text, which means […]

Where does the rigor go? Research software and the future of trustworthy science.
Generative AI now makes it dramatically easier to produce something that looks like research: analysis code, figures, literature reviews, even whole pap…

crawshaw - 2026-05-07
The industry-established code review process, review-then-commit, was a straightforward mechanism that allowed a relatively low-trust group of engineers to collaborate. It appears to have been initially developed for the Apache server OSS project in the 90s, corporatized by Google in the early 2000s, and popularized throughout the industry by several means, most notable of which was the GitHub PR.
Ten quick tips to SNIFF out sustainable and secure scientific software
Modern computational biology depends heavily on open-source software tools, analysis pipelines, and containerized workflows developed and shared by the research community. While there is extensive guidance (including Quick Tips and Simple Rules articles) on how to build robust and sustainable scientific software, far less has been written for researchers in the role of software users evaluating whether an existing tool is reliable, secure, and sustainable enough for their work. Here we present ten quick tips to help researchers critically assess the tools they adopt. Our tips are organized around a framework that centers on key evaluation features: source, network, interaction, fit, and fragility (SNIFF). These dimensions prompt researchers to consider who maintains a tool and why, whether it is embedded in a broader ecosystem, how actively its developers and users engage, whether it matches the intended use case and licensing requirements, and how robust its dependencies and security practices are. By applying these tips, researchers can make more informed decisions, reduce the risk of relying on abandoned or insecure software, and contribute to a more sustainable scientific software ecosystem.
This morning, I had the distinct honour of giving the opening keynote at #RSECon26: "Where does the rigour go? Research software and the future of trustworthy science" The core argument: AI is… | Arfon Smith
This morning, I had the distinct honour of giving the opening keynote at #RSECon26: "Where does the rigour go? Research software and the future of trustworthy science" The core argument: AI is making it dramatically cheaper to produce plausible scientific outputs, without necessarily making them cheaper to check. The scarce resources are shifting towards attention, judgment, and verification. This is no doubt a disruptive time for research software engineers and research more generally. For Research Software Engineers specifically I think that creates some important new opportunities around: encoding domain judgment, building verification infrastructure, connecting agents to reality, and helping ensure that the emerging verification layer for science remains open. Thanks to everyone at RSECon26 for the thoughtful discussion, and especially to the many people whose work and ideas I that heavily influenced this talk. Link to slides in first comment

Better tools made Copilot code review worse. Here's how we actually improved it.
How migrating Copilot code review to shared Unix-style code exploration tools reduced review cost by reshaping agent workflows around pull request evidence.

AI’s Growing Role as Scientific Peer Reviewer | Stanford HAI
Stanford computer scientist James Zou is exploring how AI can accelerate scientific research and peer review. His finding: AI excels at spotting gaps, but judgment calls still need humans.

Konrad Hinsen's blog
How can we document software and computational analyses in such a way that others can convince themselves of their validity, and build on them for their own work? The question has been around for many years, and a number of attempts have been made to provide partial answers. This post provides a brief review and describes my own tentative answer, inviting you to play with it.
Why We Can't Have Nice Software - Andrew Kelley
The problem with software is that it's too powerful. It creates so much wealth so fast that it's virtually impossible to not distribute it.
Software Should Be Allowed to Be Weird - Ewan’s Blog
A fairly large amount of the software I write would probably die in a corporate planning meeting. There would be questions about the target audience, the bu…
The Case for Software Criticism
Software may be the defining cultural artifact of our time. So why isn’t there a culture of critical analysis around it?

Sung Kim (@sungkim.bsky.social)
I like this quote on code review: "The idea that humans will reliably look through code to find the increasingly rare issues that automated tools miss seems like a fantasy." By Marc Brooker, AWS.
New preprint! We introduce a new benchmark, SciConBench, with 9.11k scientific questions derived from Cochrane Systematic Reviews. We find evidence that frontier AI agents **cannot** synthesize scientific conclusions well. A thread 🧵 w/ @hayoungjung.bsky.social & others!