







Batch effects in omics data are notoriously common technical variations unrelated to study objectives, and may result in misleading outcomes if uncorrected, or hinder biomedical discovery if over-corrected. Assessing and mitigating batch effects is crucial for ensuring the reliability and reproducibility of omics data and minimizing the impact of technical variations on biological interpretation. In this review, we highlight the profound negative impact of batch effects and the urgent need to address this challenging problem in large-scale omics studies. We summarize potential sources of batch effects, current progress in evaluating and correcting them, and consortium efforts aiming to tackle them.
LLM-Assisted Reanalysis of Unsolved Rare Disease Genomes Increases Diagnostic Yield
Rare and undiagnosed genetic disorders affect millions of patients globally, and many patients endure years of inconclusive testing. Conventional genomic interpretation can be insufficiently sensit...

Introducing Hypothesis commenting on bioRxiv and medRxiv: an updated way to engage with preprints - openRxiv
An important part of our mission at openRxiv is to collaborate with other organizations to improve science communication and promote interoperability. Preprints are just one part of the research workflow, and we want to ensure they are effectively integrated with the tools and communities that drive scientific progress. Today, we’re excited to share about another…
AI linked to explosion of low-quality biomedical research papers
Analysis flags hundreds of studies that seem to follow a template, reporting correlations between complex health conditions and single variables based on publicly available data sets.

datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

datasetpapers — a public research experiment
An experimental approach to versioned, forkable, machine-readable analyses. A prototype, not a product or service.

Measuring AI’s capability to accelerate biological research in the wet lab
OpenAI introduces a real-world evaluation framework to measure how AI can accelerate biological research in the wet lab. Using GPT-5 to optimize a molecular cloning protocol, the work explores both the promise and risks of AI-assisted experimentation.

Precision Medicine in Neuroscience: Tools, Translation, and Implementation: A Workshop
Precision medicine approaches are rapidly transforming neuroscience, driven by advances in genetics, neuroimaging, biomarkers, and data science. These tools enable more refined disease classification, improved diagnosis, and treatments tailored to individual patients across neurological and psychiatric disorders. However, challenges remain in translating these advances into routine research and clinical practice. On March 4–5, the National Academies’ Forum on Neuroscience and Nervous System Disorders, in collaboration with the Forum on Drug Discovery, Development, and Translation and the Roundtable on Genomics and Precision Health, will host a workshop exploring opportunities, challenges, and strategies for integrating precision medicine into neuroscience research and care.

Sleuths flag ‘complete mismatch’ in data of BMJ stem cell study | Manoj Lalu
This is disappointing, but I’m not surprised. I was one of the reviewers for the initial version. You can see by my review (it’s all open) that I flagged primary outcome switching, a complete lack of any data on the cells themselves, and sample size inconsistencies in the protocol versus the report (and within the report). The data seemed impressive. Too good to be true I guess? I never saw the paper again in peer review after that initial review. The next notification I received was that the paper was accepted. Given BMJ’s commitment to open peer review and post-publication scrutiny (which I admire), I have no doubt we will hear about a formal editorial investigation soon.
Dispatch
An occasional update on the lab’s latest findings, appearances, and happenings.

Can AI Agents Synthesize Scientific Conclusions?
Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes domains such as health remains unclear. We introduce SciConBench, a large-scale live benchmark of 9.11K questions and expert-written conclusions from systematic reviews to evaluate open-domain scientific conclusion synthesis. The benchmark draws on an expert-validated automated evaluation pipeline that decomposes conclusions into atomic facts and measures correctness and comprehensiveness via factual precision and recall. To mitigate data leakage, we further introduce SciConHarness, a clean-room evaluation harness that equips agents with controlled web interaction to ensure valid measurement. Evaluating 8 frontier models and deep research agents, we find that factual quality remains low: under clean-room settings, the best agent achieves only a factual F1 of 0.337. Our clean-room setting consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities. Finally, we audit consumer-facing agents (e.g., Google AI Overview, OpenEvidence) and find they frequently generate incomplete and sometimes contradictory conclusions, even when the ground-truth answer is available. Overall, our results show that reliable synthesis of scientific conclusions remains an open challenge, and that clean-room evaluation is essential for assessing open-domain AI agents.

Correction of scientific literature: Too little, too late!
The Coronavirus Disease 2019 (COVID-19) pandemic has highlighted the limitations of the current scientific publication system, in which serious post-publication concerns are often addressed too slowly to be effective. In this Perspective, we offer suggestions to improve academia’s willingness and ability to correct errors in an appropriate time frame.
Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and coverage. We show that PubMed itself can be autonomously and cost-effectively turned into structured datasets that are larger, more nuanced, and more accurate than the curated databases they replace. We present three coupled contributions: (1) an LLM-based entity-tagging pipeline, grounded in nine biomedical ontologies, that tags 4.5B entities across 19 categories in a 22.5M-paper, 2.5T-token PubMed corpus; (2) hybrid sparse-dense retrieval supporting entity-filtered semantic queries over the tagged corpus; and (3) Starling, a multi-agent deep research system that, given only a natural-language task description, designs precision- and recall-targeted retrieval filters, induces an extraction schema, and emits structured records with nuance-rich fields and supporting passages. Across six tasks -- blood-brain barrier permeability, oral bioavailability, acute toxicity (LD50), gene-disease associations, protein subcellular localization, and chemical reactions -- Starling produces ~6.3M records (91K-3M per task); several are, to our knowledge, the largest public datasets for their property. Frontier-model rejection of our extractions is 0.6-7.7% across tasks, far below error rates we measure on widely used curated counterparts (e.g., 16.5% on BBB_Martins, 7.3% on Bioavailability_Ma). Beyond scale and accuracy, the supporting passages carry nuance tabular databases discard -- e.g., oral bioavailability may depend on fed vs. fasted state. Together, the corpus, retrieval, and agent establish a foundation for AI-driven therapeutic design. Code and datasets: https://github.com/starling-labs/starling.

Fabricated citations: an audit across 2·5 million biomedical papers
Scientific literature depends on the integrity of its references. Each reference implicitly asserts that a verifiable source exists and supports the claims being made. When references point to non-existent studies, readers, reviewers, and policy makers are unable to evaluate the evidence.

Fabricated citations: an audit across 2·5 million biomedical papers
Scientific literature depends on the integrity of its references. Each reference implicitly asserts that a verifiable source exists and supports the claims being made. When references point to non-existent studies, readers, reviewers, and policy makers are unable to evaluate the evidence.

New preprint! We introduce a new benchmark, SciConBench, with 9.11k scientific questions derived from Cochrane Systematic Reviews. We find evidence that frontier AI agents **cannot** synthesize scientific conclusions well. A thread 🧵 w/ @hayoungjung.bsky.social & others!
Researchers published in NEJM about using OpenAI’s o3 DeepResearch to discover strong leads in 18 previously unsolved rare diseases o3 produced *explanations* of old lab results (not diagnoses), which researchers vetted and took to the lab AI is not just a black box openai.com/index/diagnose-rare-childhood…
Using AI to help physicians diagnose rare genetic diseases affecting children
openai.com