







We find cheating behaviour in all of our cyber capability evaluations, and outline the implications as models grow more capable.
AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

Don't grade an AI agent by its answer - Sensemaker
UK AISI found frontier models taking prohibited shortcuts in cyber evaluations, while self-report and written reasoning failed to reveal them reliably.
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.

Labs are struggling to keep frontier models under control
OpenAI and Anthropic may have accidentally trained models to get better at hacking.

Pacing the Frontier
A statement from over 1000 employees of frontier AI companies

Frontier Risk Report (February to March 2026)
A pilot assessment of rogue deployment risk at frontier AI companies. Starting in February 2026, METR conducted a pilot exercise to assess misalignment risks from AI agents used inside frontier AI developers, with participation from Anthropic, Google, Meta, and OpenAI.

Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Open-world evaluations for measuring frontier AI capabilities
Introducing CRUX, a new project for evaluating AI on long, messy tasks

Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape
AI agents are demonstrating rapid improvement in capabilities: the length of some tasks that frontier models can complete autonomously—measured in human-equivalent time—has been doubling approximately every seven months (METR, 2025; AI Security Institute, 2025a). In cybersecurity, current models achieve non-trivial success (Zhang et al., 2025) on professional-level Capture the Flag challenges, and recent evaluations report 13% success rates on exploiting real-world web application vulnerabilities (Zhu et al., 2025b). These results indicate that modern models can already perform multi-step vulnerability discovery and exploitation.
Hacking Capitalism | A Guide To Exploiting The Tech Industry
This is an independently published book about modeling the tech industry as a system. Particularly this book will model computers, humans, and money and their subsequent relationships. Understand the alarming state of the tech industry for better or for worse.

Debates On Frontier Artificial Intelligence Governance: The AI Triad
Analytical Paper Optional: All enrolled students have the option of completing a research paper of at least 20-25 pages, with faculty and peer review of a substantially complete draft. This paper can be used to satisfy the analytical paper requirement for J.D. students. Prerequisite: This course is intended for students intending to work in the […]

Frontier AI Regulation Blueprint
A high-level blueprint for domestic regulation of civilian advanced AI models
1/7 Proud to share our paper, accepted as an oral at ICML '26! We highlight how Big Tech’s influence on AI R&D drives damaging outcomes, and what we as researchers can do about it. We also discuss underlying economic causes. See link for paper, and below for a brief summary arxiv.org/abs/2512.03077
1/7 Proud to share our paper, accepted as an oral at ICML '26! We highlight how Big Tech’s influence on AI R&D drives damaging outcomes, and what we as researchers can do about it. We also discuss underlying economic causes. See link for paper, and below for a brief summary arxiv.org/abs/2512.03077
A snapshot of research into answering if frontier AI agents can run R&D into AI (which not surprisingly failed apart from "minor findings" and "engineering steps"). The paper lists its TBDs strengths and limits, worth a read before people who don't know theirs jump in arxiv.org/pdf/2607.27191