







Introducing CRUX, a new project for evaluating AI on long, messy tasks
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.
Open Source AI Policy Landscape — RedMonk Analysis
How foundations and projects are responding to AI-generated contributions. This analysis surveys 88 major organizations.
Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.

AI Research Evaluation: Negative Findings and Failure Modes | Arvind Narayanan posted on the topic | LinkedIn
📢AI agents can autonomously conduct AI research when the result is easily verifiable, but what about open-ended AI research? That’s much harder to study, and our new preprint is our first crack at doing so. Our main finding is negative, and we identify five recurring failure modes. https://lnkd.in/eGKYi4Sa Our results are tentative, and we are working to address the limitations (sample size, potential scaffold improvements). But if the finding holds up, what are the implications? It depends on whether you think recursive self improvement can be achieved simply by hill climbing at scale (I personally don’t think so) and whether you think current limitations of open-ended research like judgment and creativity could change quickly (I’m personally very open to this possibility). We plan to continue this style of evaluation — which we call shadow evaluation — on a regular basis. We’ve wanted to do this for two years, but it took so long because we wanted to get the method right. The idea behind shadow evaluation was suggested by some of the UK AISI coauthors of the paper and refined by the Princeton team. This method has important advantages (and limitations) over the current ways of evaluating agents’ ability to conduct AI research. If you’re an AI researcher interested in working with us on a shadow evaluation based on one of your papers, we’d love to hear from you. https://lnkd.in/ecyp55SW This type of evaluation necessarily involves a ton of researcher flexibility in design, execution, and interpretation. Members of the core team have a particular position in the debate on recursive self-improvement / superintelligence, and this could influence how we conduct the research. We have a detailed section in the paper on our potential biases and how we address them. We sought out a team of collaborators who don’t all share our priors, and we explicitly surface the interpretive disagreements that resulted. For future evaluations, we are interested in having “adversarial collaborators” as part of the core team. This paper exists because of the careful, time-consuming and very much human work that Peter Kirgis, Sayash Kapoor, Andrew Schwartz, and Stephan Rabanser did over the last few months. I’m also very grateful to the larger group of collaborators and co-authors. The work is part of the larger CRUX project that pushes frontier AI agents beyond what benchmarks can measure (https://cruxevals.com/). We are looking for a senior researcher to join the team: https://lnkd.in/e9dC22X5
Measuring AI Ability to Complete Long Tasks
We propose measuring AI performance in terms of the *length* of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months. Extrapolating this trend predicts that, in under a decade, we will see AI agents that can independently complete a large fraction of software tasks that currently take humans days or weeks.

Can AI agents conduct open-ended AI research? Early evidence from two case studies
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

The Open-Source Toolkit for Building AI Agents v2
An opinionated, developer-first guide to building AI agents with real-world impact

Meet Foundry: An AI Startup that Builds, Evaluates, and Improves AI Agents

Introduction to Agents
Discover what actually works in AI. Join millions of builders, researchers, and labs evaluating agents, models, and frontier technology through crowdsourced benchmarks, competitions, and hackathons.

Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.

The AI "Evaluation Crisis" Is an Opportunity to Get Data Flow Right
Why the AI evaluation crisis could force a reckoning on dataset provenance, attribution, and consent.

Built to benefit everyone
OpenAI’s recapitalization strengthens mission-focused governance, expanding resources to ensure AI benefits everyone while advancing innovation responsibly.

Debates On Frontier Artificial Intelligence Governance: The AI Triad
Analytical Paper Optional: All enrolled students have the option of completing a research paper of at least 20-25 pages, with faculty and peer review of a substantially complete draft. This paper can be used to satisfy the analytical paper requirement for J.D. students. Prerequisite: This course is intended for students intending to work in the […]

A snapshot of research into answering if frontier AI agents can run R&D into AI (which not surprisingly failed apart from "minor findings" and "engineering steps"). The paper lists its TBDs strengths and limits, worth a read before people who don't know theirs jump in arxiv.org/pdf/2607.27191

AI 2027

What will be left for us to work on?

OpenAI and Hugging Face partner to address security incident during model evaluation
Taking more seriously the claim that recent ML models do not "reason", it still is quite odd the particular ways that superhuman game-playin…

General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

StoryScope: Investigating idiosyncrasies in AI fiction