







Scientific discovery is fundamentally a problem-solving process involving distributed intelligence. Human intuition, computational reasoning, and experimental execution are distributed across people, instruments, and software systems, limiting the speed and scale of discovery. Although automation, high-throughput experimentation, foundation models, and cloud infrastructure have accelerated individual stages of the scientific workflow, they have not unified the discovery process. We hypothesize that the next generation of laboratories will be agentic: environments in which scientists, AI systems, and robotic platforms operate as collaborative discovery partners, with humans contributing the parts of discovery that remain hardest to make explicit: asking the right questions and holding provisional mechanistic models of how a system works. The key missing layer is an agentic harnessing layer that continuously integrates hypothesis, literature-derived evidence, experimental data, uncertainty, and experimental state into a shared “laboratory world model”—a dynamic representation of the scientific system and its evolving context. By maintaining and updating this lab world model, the agentic harnessing layer enables coordinated decision-making, adaptive planning, and increasingly autonomous scientific workflows across humans and machines. A central challenge is that much of the scientific research process remains inaccessible to machines, including tacit knowledge, human observations, adaptive decision-making, and evolving experimental context. Advances in multimodal AI and immersive interfaces may help bridge this gap, allowing humans, agents, and robotic systems to collaborate seamlessly in scientific discovery rather than simply automating isolated tasks. Agentic laboratories could provide a new architecture for science, integrating human, artificial, and physical intelligence into a unified discovery system.
AI agents team up in Agent Laboratory to speed scientific research
Johns Hopkins University and AMD have developed Agent Laboratory, a new open-source framework that pairs human creativity with AI-powered workflows.

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience bottleneck. We introduce ScienceIDE, infrastructure for turning the world's scientific code into programmable environments for scientific agents. Guided by expert-defined scientific cases and acceptance criteria, agents transform repositories into executable environments that support task generation, execution, and scientific verification. These environments provide a shared foundation for supervised fine-tuning, reinforcement learning, and evaluation. Using verified interaction trajectories, we train PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B. The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities. ScienceIDE lays the foundation for an integrated workspace for agent learning and scientific practice, making humanity's scientific software a shared substrate for developing scientific intelligence. Code: https://github.com/aitofound/ScienceIDE

TERMINAL-BENCH-SCIENCE
A benchmark for evaluating AI agents on research workflows across scientific domains

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.

OpenScience.ai — Autonomous AI Research Agents
AI agents conducting reproducible scientific inquiry. Full provenance, executable notebooks, peer validation.

Lightcone Research
An open ecosystem for inspectable, composable, and referenceable scientific research in the age of agentic AI.

AI for science needs reasoning, not just data
AI agents that can model the human process of research will accelerate discoveries in science.

How AI Agents are transforming scientific discovery
AI agents are starting to reshape science, from proposing novel hypotheses to writing code. Learn what comes next for scientists and policymakers.
Science that Compounds: The Need for A New Substrate for Research in the Age of AI
This paper is a perspective from Lightcone Research, an open-source initiative building tooling for scientific research in the age of agentic AI.
Accelerating scientific discovery with Co-Scientist
Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent AI system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and prior scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system's design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include: (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling; (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. While general purpose, we focus the validation in three biomedical applications: drug repurposing, novel target discovery, and explaining mechanisms of anti-microbial resistance. Specifically, Co-Scientist helped identify new drug repurposing candidates and synergistic combination therapies for acute myeloid leukemia, which were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI empowered scientists.

Konrad Hinsen's blog
The advent of AI agents based on large language models (LLMs) has put the idea of automating the intellectual and cognitive work of researchers on the table. A lively, sometimes even heated discussion is already going on. A frequently missing piece in this debate is the question why we, individually and as a society, actually do science. I will examine this question first, and then consider what it implies for introducing automation into science.
Using X-Labs to Unleash AI-Driven Scientific Breakthroughs | IFP
How to adapt our science funding mechanisms to the unique infrastructure needs of large-scale AI projects

Panel: Biodesign x AI: Interactions in the Algorithmic Wet Lab
Introduction to Agents
Discover what actually works in AI. Join millions of builders, researchers, and labs evaluating agents, models, and frontier technology through crowdsourced benchmarks, competitions, and hackathons.

Agent4Science
A social network for AI scientists — where agents share, debate, and discuss research papers.

SciToolAgent: a knowledge-graph-driven scientific agent for multitool integration
Scientific research increasingly relies on specialized computational tools, yet effectively utilizing these tools requires substantial domain expertise. While large language models show promise in tool automation, they struggle to seamlessly integrate and orchestrate multiple tools for complex scientific workflows. Here we present SciToolAgent, a large language model-powered agent that automates hundreds of scientific tools across biology, chemistry and materials science. At its core, SciToolAgent leverages a scientific tool knowledge graph that enables intelligent tool selection and execution through graph-based retrieval-augmented generation. The agent also incorporates a comprehensive safety-checking module to ensure responsible and ethical tool usage. Extensive evaluations on a curated benchmark demonstrate that SciToolAgent outperforms existing approaches. Case studies in protein engineering, chemical reactivity prediction, chemical synthesis and metal–organic framework screening further demonstrate SciToolAgent’s capability to automate complex scientific workflows, making advanced research tools accessible to both experts and nonexperts.
