







ABSTRACT Large language models (LLMs) like ChatGPT seem to be increasingly used for information seeking and analysis, including to support academic literature reviews. To test whether the results might sometimes include retracted research, we identified 217 retracted or otherwise concerning academic studies with high altmetric scores and asked ChatGPT 4o‐mini to evaluate their quality 30 times each. Surprisingly, none of its 6510 reports mentioned that the articles were retracted or had relevant errors, and it gave 190 relatively high scores (world leading, internationally excellent, or close). The 27 articles with the lowest scores were mostly accused of being weak, although the topic (but not the article) was described as controversial in five cases (e.g., about hydroxychloroquine for COVID‐19). In a follow‐up investigation, 61 claims were extracted from retracted articles from the set, and ChatGPT 4o‐mini was asked 10 times whether each was true. It gave a definitive yes or a positive response two‐thirds of the time, including for at least one statement that had been shown to be false over a decade ago. The results therefore emphasise, from an academic knowledge perspective, the importance of verifying information from LLMs when using them for information seeking or analysis.
Does <span style="font-variant:small-caps;">ChatGPT</span> Ignore Article Retractions and Other Reliability Concerns?
ABSTRACT Large language models (LLMs) like ChatGPT seem to be increasingly used for information seeking and analysis, including to support academic literature reviews. To test whether the results might sometimes include retracted research, we identified 217 retracted or otherwise concerning academic studies with high altmetric scores and asked ChatGPT 4o‐mini to evaluate their quality 30 times each. Surprisingly, none of its 6510 reports mentioned that the articles were retracted or had relevant errors, and it gave 190 relatively high scores (world leading, internationally excellent, or close). The 27 articles with the lowest scores were mostly accused of being weak, although the topic (but not the article) was described as controversial in five cases (e.g., about hydroxychloroquine for COVID‐19). In a follow‐up investigation, 61 claims were extracted from retracted articles from the set, and ChatGPT 4o‐mini was asked 10 times whether each was true. It gave a definitive yes or a positive response two‐thirds of the time, including for at least one statement that had been shown to be false over a decade ago. The results therefore emphasise, from an academic knowledge perspective, the importance of verifying information from LLMs when using them for information seeking or analysis.

Papers and peer reviews with evidence of ChatGPT writing
Retraction Watch readers have likely heard about papers showing evidence that they were written by ChatGPT, including one that went viral. We and others have reported on the phenomenon. Here’…

Influential study touting ChatGPT in education retracted over red flags
The retracted study on ChatGPT in education was already cited hundreds of times.

I saw this meta-analysis shared a couple of times recently, so we took a look and re-analyzed the data. | František Bartoš
I saw this meta-analysis shared a couple of times recently, so we took a look and re-analyzed the data. We found that the conclusion is almost entirely driven by publication bias. Both state-of-the-art and standard methods reduce the degree of the effect 2-3 fold. Moreover, the data no longer show statistical evidence for the main conclusions. Importantly, our findings do not imply that there is no positive effect of ChatGPT, or other large language models on learning; in fact, our analysis reveals that there is “absence of evidence” rather than “evidence of absence”. The present literature appears to be contaminated by publication bias; high-quality registered reports are needed to properly evaluate the effect of large language models in educational settings. See the full response just submitted for publication at https://lnkd.in/eGKHHKtg
More than 10,000 research papers were retracted in 2023 — a new record
The number of articles being retracted rose sharply this year. Integrity experts say that this is only the tip of the iceberg.

More than 10,000 research papers were retracted in 2023 — a new record
The number of articles being retracted rose sharply this year. Integrity experts say that this is only the tip of the iceberg.

The State of Papers, Retractions, and Preprints: Evidence from the CrossRef Database (2004-2024)
A 20-year analysis of CrossRef metadata demonstrates that global scholarly output -- encompassing publications, retractions, and preprints -- exhibits strikingly inertial growth, well-described by exponential, quadratic, and logistic models with nearly indistinguishable goodness-of-fit. Retraction dynamics, in particular, remain stable and minimally affected by the COVID-19 shock, which contributed less than 1% to total notices. Since 2004, publications doubled every 9.8 years, retractions every 11.4 years, and preprints at the fastest rate, every 5.6 years. The findings underscore a system primed for ongoing stress at unchanged structural bottlenecks. Although model forecasts diverge beyond 2024, the evidence suggests that the future trajectory of scholarly communication will be determined by persistent systemic inertia rather than episodic disruptions -- unless intentionally redirected by policy or AI-driven reform.

The associations of social media attention, visibility, disinformation and retraction initiators with time to retraction: a Cox regression analysis
Purpose This study examines how retraction reasons, retraction initiators, journal visibility, access models and Twitter activity associate with the speed of retracting flawed scientific publications. Design/methodology/approach Using a Cox proportional hazards model, we analyzed 1,179 articles retracted in 2019–2021, including a subset of 98 papers tweeted before retraction. Findings The results reveal that higher journal impact factor and open-access status were associated with faster retractions. However, a significant negative interaction indicated that the effect of high-impact journals diminished for open-access publications. Journal-initiated retractions were slower overall, except in cases of misconduct such as co-authorship deception and plagiarism, where journals acted more quickly. Among retraction reasons, only deception in co-authoring was associated with significantly slower retractions, but this trend reversed when journals led the process. The association of social media attention with retraction speed was statistically robust, albeit modest in magnitude: each additional pre-retraction tweet was associated with a slight reduction in time to retraction. Bootstrap validation confirmed the stability of this finding. Originality/value Public scrutiny, institutional responsibility and publication visibility jointly shape the time to retraction. This study advances altmetrics discourse by positioning social media as a conditional, yet meaningful, participant in the retraction lifecycle. Beyond altmetrics, our findings highlight retractions as part of a broader network of relationships between public accountability, digital ethics and science communication, positioning them as moments of accountability shaped jointly by journals, ethical responsibilities and digital publics.

Correction of scientific literature: Too little, too late!
The Coronavirus Disease 2019 (COVID-19) pandemic has highlighted the limitations of the current scientific publication system, in which serious post-publication concerns are often addressed too slowly to be effective. In this Perspective, we offer suggestions to improve academia’s willingness and ability to correct errors in an appropriate time frame.
How Ten Publishers Retract Research
Retractions are the primary mechanism for correcting the scholarly record, yet publishers differ markedly in how they use them. We present a bibliometric analysis of 46,087 retractions across 10 major publishers using data from the Retraction Watch database (1997-2026), examining retraction rates, reasons, temporal trends, and geographic distributions, among other dimensions. Normalized retraction rates vary by two orders of magnitude, from Elsevier's 3.97 per 10,000 publications to Hindawi's 320.02. China-affiliated authors account for the largest share of retractions at every publisher. Retraction lags and reason profiles also vary widely across publishers. Among the ten publishers, ACM is an outlier in its retraction profile. ACM's normalized rate is mid-range (5.65), yet 98.3% of its 354 retractions are related to one incident. Seven of the ten most common global retraction reasons (including misconduct, plagiarism, and data concerns) are entirely absent from ACM's record. ACM's first retraction dates to 2020, despite a catalog dating to 1997. ACM self-describes its retraction threshold as "extremely high." We discuss this threshold in relation to the COPE retraction guidelines and the implications of ACM's non-public dark archive of removed works.

Scientific production in the era of Large Language Models
Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.

Editor’s Note: Retraction of article containing fabricated quotations
We are reinforcing our editorial standards following this incident.

Post-publication critique at top-ranked journals across scientific disciplines: a cross-sectional assessment of policies and practice
Journals exert considerable control over letters, commentaries and online comments that criticize prior research (post-publication critique). We assessed policies (Study One) and practice (Study Two) related to post-publication critique at 15 top-ranked journals in each of 22 scientific disciplines ( N = 330 journals). Two-hundred and seven (63%) journals accepted post-publication critique and often imposed limits on length (median 1000, interquartile range (IQR) 500–1200 words) and time-to-submit (median 12, IQR 4–26 weeks). The most restrictive limits were 175 words and two weeks; some policies imposed no limits. Of 2066 randomly sampled research articles published in 2018 by journals accepting post-publication critique, 39 (1.9%, 95% confidence interval [1.4, 2.6]) were linked to at least one post-publication critique (there were 58 post-publication critiques in total). Of the 58 post-publication critiques, 44 received an author reply, of which 41 asserted that original conclusions were unchanged. Clinical Medicine had the most active culture of post-publication critique: all journals accepted post-publication critique and published the most post-publication critique overall, but also imposed the strictest limits on length (median 400, IQR 400–550 words) and time-to-submit (median 4, IQR 4–6 weeks). Our findings suggest that top-ranked academic journals often pose serious barriers to the cultivation, documentation and dissemination of post-publication critique.

'Nature' Retracts Paper on the Benefits of ChatGPT in Education
“What educators, parents and policy officials really needed was high quality data and evidence to help guide them. What they have had to deal with instead is some substandard research.”
A recent experience with ChatGPT 5.5 Pro
We are all having to keep revising upwards our assessments of the mathematical capabilities of large language models. I have just made a fairly large revision as a result of ChatGPT 5.5 Pro, to whi…

An article about data visualization was retracted 1.5 years after I pointed out errors. The notice says that "concerns were raised". I spend dozens of hours contacting authors and editors, reproducing analyses, and following up on ignored emails. But I'm not mentioned in the retraction notice.
RETRACTED: A Perception Study for Unit Charts in the Context of Large-Magnitude Data Representation
www.mdpi.com