







Tens of thousands of publications from 2025 might include invalid references generated by AI, a Nature analysis suggests.
Hallucinated citations are polluting the scientific literature. What can be done?
Tens of thousands of publications from 2025 might include invalid references generated by AI, a Nature analysis suggests.

Hallucinated citations are polluting the scientific literature. What can be done?
Tens of thousands of publications from 2025 might include invalid references generated by AI, a Nature analysis suggests.

Fraudulent citations, blamed on AI hallucinations, are becoming more common in research papers
“Fabricated” citations that do not reference real academic papers are spreading in the literature, polluting the public record of science, a new study found

Fabricated citations: an audit across 2·5 million biomedical papers
Scientific literature depends on the integrity of its references. Each reference implicitly asserts that a verifiable source exists and supports the claims being made. When references point to non-existent studies, readers, reviewers, and policy makers are unable to evaluate the evidence.

Fabricated citations: an audit across 2·5 million biomedical papers
Scientific literature depends on the integrity of its references. Each reference implicitly asserts that a verifiable source exists and supports the claims being made. When references point to non-existent studies, readers, reviewers, and policy makers are unable to evaluate the evidence.

The Discovery Engine: A Framework for AI-Driven Synthesis and Navigation of Scientific Knowledge Landscapes
Scientific progress relies on the effective accumulation, synthesis, and critical evaluation of knowledge. Traditionally, the well-documented, peer reviewed publication served as the primary standard for filtering and disseminating credible findings within the scientific community. Recently, however, we are witnessing an unprecedented acceleration in research output, a veritable explosion of scientific publications across all disciplines [1]. Yet, this very abundance creates a paradox: the sheer volume threatens to overwhelm the mechanisms designed for its assimilation and synthesis. Researchers, even within highly specialized subfields, face an almost insurmountable challenge in keeping abreast of relevant developments, integrating disparate findings, and identifying the truly novel signals amidst the noise [2]. This information overload contributes to disciplinary fragmentation, hindering the cross-pollination of ideas essential for disruptive innovation [3]. Furthermore, persistent concerns regarding "reproducibility crisis" [2], predatory journals, inflation of research areas[4], growing retractions and the potential influences of bibliometrics on research direction [5] highlight systemic challenges in validating and prioritizing scientific contributions to fundamental knowledge.
Incorrect Citation Association for Articles in Online-Only Springer Nature Journals
We show that citation metrics of journal articles in many of the online-only Springer Nature journals and associated ones are distorted, going back to articles from 2001. We find that most likely due to an API response error, there are many incorrect references which typically lead to Article Number 1 of a given Volume. Among others, the issue affects journals such as Scientific Reports, Nature Communications, Communications journals, Cell Death & Disease, Light: Science & Applications, as well as many BMC, Discovery and npj journals. Beyond the negative effect of introducing incorrect reference information, this distorts the citation statistics of articles in these journals, with a few articles being massively over-cited compared to their peers, while many lose citations; e.g. both in Scientific Reports and in Nature Communications, 5 of the 10 top cited articles have article numbers of 1. We validate the distorted statistics by assessing data from multiple scientific literature databases: Crossref, OpenCitations, Semantic Scholar, and the journals' websites. The issue primarily arises from the inconsistent transition from page-based referencing of articles to article number-based referencing, as well as the improper handling of the change in the publisher's article metadata API. It seems that the most pressing problem has been present since approximately 2011, which we estimate affects the citation count of millions of authors.

Quotation errors in general science journals
Due to the incremental nature of scientific discovery, scientific writing requires extensive referencing to the writings of others. The accuracy of this referencing is vital, yet errors do occur. These errors are called ‘quotation errors’. This paper presents the first assessment of quotation errors in high-impact general science journals. A total of 250 random citations were examined. The propositions being cited were compared with the referenced materials to verify whether the propositions could be substantiated by those materials. The study found a total error rate of 25%. This result tracks well with error rates found in similar studies in other academic fields. Additionally, several suggestions are offered that may help to decrease these errors and make similar studies more feasible in the future.

Fabricated Citations — CITADEL
4,046 fabricated citations in the biomedical literature. A more than 12× rise in three years. Verified across PubMed, Crossref, OpenAlex, Google Scholar.
Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection
Large language models (LLMs) are rapidly being adopted as research assistants, particularly for literature review and reference recommendation, yet little is known about whether they introduce demographic bias into citation workflows. This study systematically investigates gender bias in LLM-driven reference selection using controlled experiments with pseudonymous author names. We evaluate several LLMs (GPT-4o, GPT-4o-mini, Claude Sonnet, and Claude Haiku) by varying gender composition within candidate reference pools and analyzing selection patterns across fields. Our results reveal two forms of bias: a persistent preference for male-authored references and a majority-group bias that favors whichever gender is more prevalent in the candidate pool. These biases are amplified in larger candidate pools and only modestly attenuated by prompt-based mitigation strategies. Field-level analysis indicates that bias magnitude varies across scientific domains, with social sciences showing the least bias. Our findings indicate that LLMs can reinforce or exacerbate existing gender imbalances in scholarly recognition. Effective mitigation strategies are needed to avoid perpetuating existing gender disparities in scientific citation practices before integrating LLMs into high-stakes academic workflows.

Okay so, we just found that over 50 papers published at @Neurips 2025 have AI hallucinations by @alexcdot(Alex Cui) | Twitter Thread Reader
Okay so, we just found that over 50 papers published at @Neurips 2025 have AI hallucinations I don't think people realize how bad the slop is right now It's not just that researchers from @GoogleDeepMind, @Meta, @MIT, @Cambridge_Uni are using AI - they allowed LLMs to generate hallucinations in their papers and didn't notice at all. It's insane that these made it through peer review👇

White House Health Report Included Fake Citations (Published 2025)
A report on children’s health released by the Make America Healthy Again Commission referred to scientific papers that did not exist.

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

Halupedia: An AI-Generated Wikipedia-Style Encyclopedia of Fabricated Knowledge and Absurd AI Fabulation - BizTech Weekly
Analysis of Halupedia’s AI-driven on-demand encyclopedia model reveals real-time, non-persistent article generation that simulates authoritative references through fabricated citations and internal “canon” consistency, highlighting challenges in provenance, hallucination, moderation, and the evolving trade-offs between novelty-driven engagement and information integrity in generative AI systems.

Papers and patents are becoming less disruptive over time
Theories of scientific and technological change view discovery and invention as endogenous processes1,2, wherein previous accumulated knowledge enables future progress by allowing researchers to, in Newton’s words, ‘stand on the shoulders of giants’3–7. Recent decades have witnessed exponential growth in the volume of new scientific and technological knowledge, thereby creating conditions that should be ripe for major advances8,9. Yet contrary to this view, studies suggest that progress is slowing in several major fields10,11. Here, we analyse these claims at scale across six decades, using data on 45 million papers and 3.9 million patents from six large-scale datasets, together with a new quantitative metric—the CD index12—that characterizes how papers and patents change networks of citations in science and technology. We find that papers and patents are increasingly less likely to break with the past in ways that push science and technology in new directions. This pattern holds universally across fields and is robust across multiple different citation- and text-based metrics1,13–17. Subsequently, we link this decline in disruptiveness to a narrowing in the use of previous knowledge, allowing us to reconcile the patterns we observe with the ‘shoulders of giants’ view. We find that the observed declines are unlikely to be driven by changes in the quality of published science, citation practices or field-specific factors. Overall, our results suggest that slowing rates of disruption may reflect a fundamental shift in the nature of science and technology.

AI agents are checking the scientific literature — and spotting decades-old errors
The technology is proving adept at finding faults in decades-old papers and reference databases.
