







Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers. He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them.
Transactions on Machine Learning Research
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission? medium.com/@TmlrOrg/asking-authors-about…
Sep 16, 2026 at 8:36 PM
Gautam Kamath on Twitter / X
Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers. He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them. https://t.co/6CJF152u7I— Gautam Kamath (@thegautamkamath) September 16, 2026
Thomas G. Dietterich on Twitter / X
We are seeing a new trend in submissions to @arxiv (and presumably to conferences and journals): Authors submitting papers whose contents they likely do not understand. 1/— Thomas G. Dietterich (@tdietterich) September 13, 2026
Okay so, we just found that over 50 papers published at @Neurips 2025 have AI hallucinations by @alexcdot(Alex Cui) | Twitter Thread Reader
Okay so, we just found that over 50 papers published at @Neurips 2025 have AI hallucinations I don't think people realize how bad the slop is right now It's not just that researchers from @GoogleDeepMind, @Meta, @MIT, @Cambridge_Uni are using AI - they allowed LLMs to generate hallucinations in their papers and didn't notice at all. It's insane that these made it through peer review👇

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
Eric Topol on Twitter / X
This is a FRAUDULENT paper, AI-generated. My name was used as an author and I had nothing to do with it, never saw it until today https://t.co/Ky60zJrMEZThe "Editors" Angelo Rossi Mori, David Mensah, and Zarnie Khadjesari should be reported. pic.twitter.com/2Z5CE8w4bn— Eric Topol (@EricTopol) April 22, 2026
What it means that the AI research community can't quit twitter - rl-blogging
Some quickly jotted thoughts working through the implications
Shreya Shankar on Twitter / X
This problem has gotten significantly worse. As Twitter and other social media have become primary channels for sharing research, academics are now expected to make the same ideas legible and appealing to both the general public (to go "viral") and senior scholars (to get the… https://t.co/zPF3tKzbVa— Shreya Shankar (@sh_reya) July 25, 2026
Cas (Stephen Casper) on Twitter / X
It is hard to overstate how disappointing I think this new paper from Oxford, OpenAI, Anthropic, and Google (et al) is. I can't take it seriously as academic work, just as propaganda. It also has some very bad scholarship and questionable adherence to research ethics. Having… pic.twitter.com/Z5fBx360ya— Cas (Stephen Casper) (@StephenLCasper) May 13, 2026

Noam Brown on Twitter / X
And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.— Noam Brown (@polynoamial) August 1, 2026
Adil on Twitter / X
being assigned to review AI generated slop papers is very frustrating. the way this is supposed to work is that first you take the time to write something and then i take the time to read it. if you don't do the first part then i shouldn't have to do the second part.— Adil (@adilsoubki) August 26, 2026
Joao Pereira's Twitter Thread | Xunroll
I do not understand why we are all piling on one paper. Is there anything obviously wrong with it, aside from personal views on length? Any data anomalies ...

Publishing your work increases your luck
In 12 months, @aarondfrancis changed his life by bypassing fear and embracing risk. Now, he’s working his dream job @tuple. Get his full story on The ReadME Project:

Q&A from the slop trenches – GeoSpatial ML
We reviewed 22 ML conference submissions this summer. Fifteen had fabricated citations, hallucinated authors, or clear LLM slop, so we complain about that a bit, and release the paper references audit we now run.
How times change! A decade ago, a failed replication of work by the same lab led to 'repligate' and blog posts such as psychol.cam.ac.uk/cece/blog Now, failed replications can be published an no one blinks an eye. That is progress due to all the scholars working hard to improve science!
at this point no one cares, but i was invited to talk about the high-rep retraction saga, and once again, i find myself perplexed at the exclusion of confirmatory results in any calculation in a study making a key claim about the inclusion of confirmatory studies making results highly replicable.