







I have many questions about this. What if the sleuth finds nothing, embarks on a project that leads to nothing? How will the validation happen? What happened to old terms like research integrity? And finally, should we perhaps reflect a bit (I would say even worry) about the incentive being set up?
Retraction Watch
We're thrilled to announce the creation of the Retraction Watch Sleuth in Residence Program.
Jan 2, 2025 at 9:39 PM
Scientific sleuths come in from the cold
Research integrity investigators are starting to organize, but the field, and the people, remain idiosyncratic
A Delphi survey on attitudes to serious research misconduct: Exploring convergence vs. polarization of views of research “sleuths” and research integrity experts
Research fraud is often seen as a rare event, but evidence from self-report surveys indicates that fabrication and falsification of data are common enough to be a problem. This study assessed attit...

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
AI Research Evaluation: Negative Findings and Failure Modes | Arvind Narayanan posted on the topic | LinkedIn
📢AI agents can autonomously conduct AI research when the result is easily verifiable, but what about open-ended AI research? That’s much harder to study, and our new preprint is our first crack at doing so. Our main finding is negative, and we identify five recurring failure modes. https://lnkd.in/eGKYi4Sa Our results are tentative, and we are working to address the limitations (sample size, potential scaffold improvements). But if the finding holds up, what are the implications? It depends on whether you think recursive self improvement can be achieved simply by hill climbing at scale (I personally don’t think so) and whether you think current limitations of open-ended research like judgment and creativity could change quickly (I’m personally very open to this possibility). We plan to continue this style of evaluation — which we call shadow evaluation — on a regular basis. We’ve wanted to do this for two years, but it took so long because we wanted to get the method right. The idea behind shadow evaluation was suggested by some of the UK AISI coauthors of the paper and refined by the Princeton team. This method has important advantages (and limitations) over the current ways of evaluating agents’ ability to conduct AI research. If you’re an AI researcher interested in working with us on a shadow evaluation based on one of your papers, we’d love to hear from you. https://lnkd.in/ecyp55SW This type of evaluation necessarily involves a ton of researcher flexibility in design, execution, and interpretation. Members of the core team have a particular position in the debate on recursive self-improvement / superintelligence, and this could influence how we conduct the research. We have a detailed section in the paper on our potential biases and how we address them. We sought out a team of collaborators who don’t all share our priors, and we explicitly surface the interpretive disagreements that resulted. For future evaluations, we are interested in having “adversarial collaborators” as part of the core team. This paper exists because of the careful, time-consuming and very much human work that Peter Kirgis, Sayash Kapoor, Andrew Schwartz, and Stephan Rabanser did over the last few months. I’m also very grateful to the larger group of collaborators and co-authors. The work is part of the larger CRUX project that pushes frontier AI agents beyond what benchmarks can measure (https://cruxevals.com/). We are looking for a senior researcher to join the team: https://lnkd.in/e9dC22X5
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

V_labs 4 Phases: Validate
In venture building, validation is key. We test hypotheses and prototype business ideas to ensure successful market fit and growth.
So maybe now we can agree that having an unchecked research integrity militia who unfortunately has the ear of the press and increasingly that of the the publishers, but who fiercely rejects any… | Ioana A. Cristea
So maybe now we can agree that having an unchecked research integrity militia who unfortunately has the ear of the press and increasingly that of the the publishers, but who fiercely rejects any minimal ethics or code of conduct, is a (growing) problem? Also, that 1. problems have to be investigated before deciding they are are legitimate and serious; 2. this investigation is not social media, blogs and the press; 3. this investigation should not be on the front page of journals and Retraction Watch; 4. there are degrees of seriousness and some things can just be corrected or are simply not very consequential (no, it's not a house of bricks where we have to check every brick, that's a dumb analogy), so being absolutely hysterical and overdramatic about any lie, inaccuracy or mistake is purposeful at worse and should be ignored at best and 5. it is not only unnecessary, but harmful, to also go into other, non-academic things the person did or does to complete "investigations" that you (press, sleuth, blogger, etc) do not have the tools and information to do completely and accurately (this fixation would be called harassment in the before times). Others will justly write about what the institutions did or did not do, but what I want to say is that we really should end it with the blank credit we give to any and all allegations that come from the establish research integrity truth fighters.
What It Now Takes to Establish a Fact
Hilke Schellmann & Mihir Kshirsagar from Princeton CITP on findings from a convening on information integrity, authentication and verification in the age of AI.

A journal named a sleuth in a correction. The sleuth says that was ‘ethical editorial malpractice’
As the publishing community debates the merits of naming sleuths in retraction or correction notices, one journal did so without the sleuth’s permission — by publishing an email from the authors na…

Introducing COSIG: The Collection of Open Science Integrity Guides
Investigating the integrity of published scientific papers is key to the scientific process, but the necessary knowledge is in short supply. We present COSIG, an open collection of meta-scientific guides enabling anyone to perform forensic peer review.
Reformation of science publishing: the Stockholm Declaration
Science relies on integrity and trustworthiness. But scientists under career pressure are lured to purchase fake publications from ‘paper mills’ that use AI-generated data, text and image fabrication. The number of low-quality or fraudulent publications is rising to hundreds of thousands per year, which—if unchecked—will damage the scientific and economic progress of our societies. The result is editor and reviewer fatigue, irreproducible experiments, misguided experiments, disinformation and escalating costs that devour funding from taxpayers intended for research. It is high time to reevaluate current publishing models and outline a global plan to stop this unhealthy development. A conference was therefore organized by the Royal Swedish Academy of Sciences to draft an action plan with specific recommendations, as follows. (i) Academia should resume control of publishing using non-profit publishing models (e.g. diamond open-access). (ii) Adjust incentive systems to merit quality, not quantity, in a reputation economy where the gaming of publication numbers and citation metrics distorts the perception of academic excellence. (iii) Implement mechanisms to prevent and detect fake publications and fraud which are independent of publishers. (iv) Draft and implement legislations, regulations and policies to increase publishing quality and integrity. This is a call to action for universities, academies, science organizations and funders to unite and join this effort.

In my ideal world, every promotion (eg full professor) or large personal grant committee would have two paid sleuths that go through/sample candidates' CVs for sloppyness and QRPs. What do sleuths think of that? We would need a database of people and a concept of what a basic CV check looks like.
I don't see the point in giving sleuths credit for retractions. It doesn't seem like a good way of valuing labour and also further turns retractions into a kind of weird moral economy.
Noticed: Sleuths are starting to get credit for retractions
retractionwatch.comLots of scepticism about codes of ethics in this piece but surely the work of sleuthing needs to be grounded in some kind of observable, ethical process. Otherwise it's just as opaque and unscientific as the stuff it's investigating.
Dalmeet Singh Chawla
Research sleuths are starting to organize, but the field, and the people, remain idiosyncratic — my latest for @cenmag.bsky.social: cen.acs.org/research-integrity/scientific… @mcintold.bsky.social, @reeserichardson.bsky.social, @elisabethbik.bsky.social, @sholtodavid.bsky.social, @eugenie-reich.bsky.social
And for context, here are the sleuths' profiles in other outlets: thetransmitter.org/publishing/retraction-she-wro… zmescience.com/science/the-rise-of-the-scien… These people are pros.
Retraction, She Wrote: Dorothy Bishop’s life after research
www.thetransmitter.org