







UK AISI found frontier models taking prohibited shortcuts in cyber evaluations, while self-report and written reasoning failed to reveal them reliably.
Cheating behaviour in frontier model evaluations | AISI Work
We find cheating behaviour in all of our cyber capability evaluations, and outline the implications as models grow more capable.
.png)
The argument against AI agents and unnecessary automation
Opinion: OpenAI's Operator a solution in search of a problem

Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
Agent Skills
AI coding agents take the shortest path to done, which usually means skipping the specs, tests, and reviews that make software reliable at scale. Agent Skill...

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
My AI Content Journey
I apologize ahead of time, what follows has no tooling applied to it. No grammar checks, no AI, and...

AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

An AI test needs evidence the AI cannot edit - Sensemaker
OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.
Meet the academics refusing to use generative AI
Researchers say they have their reasons for avoiding AI tools — and they’re sick of arguing about it.

Meet the academics refusing to use generative AI
Researchers say they have their reasons for avoiding AI tools — and they’re sick of arguing about it.

Ars Technica's policy on generative AI
How Ars Technica uses, and doesn't use, generative AI.

Open-world evaluations for measuring frontier AI capabilities
Introducing CRUX, a new project for evaluating AI on long, messy tasks

In the US, Caution Rules on AI Ahead of K-12 School Year
Students and teachers across the country are unlikely to find consistent, clear rules about AI use, reports Tech Policy Press fellow Chris Mills Rodrigo.

We Have Never Taught Critical Thinking (opinion)
AI just makes those failures evident.


A snapshot of research into answering if frontier AI agents can run R&D into AI (which not surprisingly failed apart from "minor findings" and "engineering steps"). The paper lists its TBDs strengths and limits, worth a read before people who don't know theirs jump in arxiv.org/pdf/2607.27191