







An author-facing AI plugin that helps you prepare a cleaner, reproducible replication package before submission.
Keynote: Reproducibility and replicability of computer simulations | Canal U
Since the early days of the reproducibility crisis, much progress has been made in understanding and improving computational reproducibility and replicability (R and R)...

Getting Started with ML and AI in Research Software | Software Sustainability Institute
Getting started with ML in research software means embracing a shift in how results are produced and reproduced. Instead of a fixed execution path, research software teams work with systems whose behaviour emerges from data, configuration, and training dynamics. Reproducibility becomes a matter of capturing the process rather than relying solely on the code. The tools and techniques outlined here can be adopted incrementally into existing projects, and together they provide a practical foundation for reproducible ML research.
Git AI - Track AI Code all the way to production
Cross-agent observability from prompt to production. Track AI-generated code from Cursor, Claude Code, GitHub Copilot, Gemini, and more through the entire SDLC.
Dialog DB "Serverless" Replication Demo
Agent Plugins
A portable package format for reusable components that extend AI agents.

GitHub - Responsible-Dataset-Sharing/easy-dataset-share: A CLI tool that helps AI researchers share datasets responsibly.
A CLI tool that helps AI researchers share datasets responsibly. - Responsible-Dataset-Sharing/easy-dataset-share
Managing Changes from Multiple AI Agents
In this video, I walk you through managing changes from multiple AI agents in our to-do app, specifically focusing on the Expert to CSV button. We explore design options using ground agents, where one agent proposes a green design and the other a blue one. After some time, both agents complete their tasks, and their changes are synchronized to GitHub and our local machine. I demonstrate the final designs and decide to ship the green version as the main one. I encourage you to apply this approach in your own projects and consider using rebase in GitHub for similar tasks.
The Last Human-Written Paper: Agent-Native Research Artifacts
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities.

Antithesis: autonomous software testing
Try the Antithesis autonomous testing platform and find bugs in your software with perfect reproducibility

GRN · German Reproducibility Network
Working together for trustworthy and useful research
Software Is Made Between Commits
From the Zed Blog: Agents turned the conversation into the real source of our software. DeltaDB is the version control built for it.
The website that created an AI clone of its editor in chief
Every CEO Dan Shipper on doubling headcount while automating everything, building an agent out of 30,000 copyedits, and the “dirty secret” of writing with AI

Reproducible Execution Environment (REE) | Tech | Gensyn
Run AI model inference in a machine-agnostic environment where the same model and inputs produce the same outputs across supported hardware.

DeepWiki | AI documentation you can talk to, for every repo
DeepWiki provides up-to-date documentation you can talk to, for every repo in the world. Think Deep Research for GitHub - powered by Devin.

🚨 AI News | TestingCatalog on Twitter / X
OPENAI 🚨: Early look at an upcoming Agent Studio for building and hosting configurable, always-on 24/7 agents on ChatGPT. The "Hermes" feature is tightly coupled with the existing Workflows builder on the OpenAI Platform, and you will also be able to open your Agents in the… https://t.co/7U4m2vNS3z pic.twitter.com/tfPIGdXRAs— 🚨 AI News | TestingCatalog (@testingcatalog) April 21, 2026