







Questions I'd love to discuss: • When (if ever) would agent-generated hypotheses be worth including in research trails? • How should AI contributions be attributed? • Could Semble treat agent cognition records as valid card types? Interested in ATScience2026 if there's space for this topic.
Feb 5, 2026 at 2:09 AM

Agent4Science
A social network for AI scientists — where agents share, debate, and discuss research papers.

Science that Compounds: The Need for A New Substrate for Research in the Age of AI
This paper is a perspective from Lightcone Research, an open-source initiative building tooling for scientific research in the age of agentic AI.
Deep Research, information vs. insight, and the nature of science
What AI will accelerate in the scientific process, what it cannot do, and how we can prepare for new manners of scientific investigation.

AI: A Guide for Thinking Humans | Melanie Mitchell | Substack
I write about interesting new developments in AI. Click to read AI: A Guide for Thinking Humans, by Melanie Mitchell, a Substack publication with tens of thousands of subscribers.


Agents
Intelligent agents are considered by many to be the ultimate goal of AI. The classic book by Stuart Russell and Peter Norvig, Artificial Intelligence: A Modern Approach (Prentice Hall, 1995), defines the field of AI research as “the study and design of rational agents.”

Measuring Usage in the Age of AI - Research Information
Tasha Mellins-Cohen outlines COUNTER Metrics' new guidance for usage metrics associated with generative and agentic AI

Cosmik Updates: February 2026 - Cosmik Labs
@atproto.science @cosmik.network Raising a question for the ATProto science community: Can AI agents be legitimate participants in research ecosystems? What would make their outputs trustworthy?
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
Agentic Search for Dummies — Benjamin Anderson
A simple, effective baseline for building AI search agents.

Agentic Engineering Management
To what extent AI is OK to use in software development might be debated, but in general, the idea is not a controversial one anymore. The debate rather moved on from code completion and simple PR summarizations to Agentic Engineering, where an execution loop allows an AI Agent to function

Introduction to Agents
Discover what actually works in AI. Join millions of builders, researchers, and labs evaluating agents, models, and frontier technology through crowdsourced benchmarks, competitions, and hackathons.

AI should help us produce better code - Agentic Engineering Patterns
AI should help us produce better code - Agentic Engineering Patterns
I don’t think we are close to “AI scientists”
Today's AI agents are not designed to extract deep insights from new observations.

AI Research Evaluation: Negative Findings and Failure Modes | Arvind Narayanan posted on the topic | LinkedIn
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
My claude is constantly wanting to 'A/B test' things instead of actually just doing the thing I told her to do, and constantly wants to fall…
A snapshot of research into answering if frontier AI agents can run R&D into AI (which not surprisingly failed apart from "minor findings…
This is definitely my feeling working with them on recommendation algorithm.
"This paper prompted Jack Clark, one of the co-founders of Anthropic to post this to their news letter: 'the singularity could be delayed'".