







View and interact with ARC-AGI task 39a8645d
ARC-AGI-3
ARC-AGI-3 is the first interactive reasoning benchmark for AI agents—play as humans and build agents that learn in novel environments.

Alex MacCaw on Twitter / X
I suspect generalized reasoning was solved just a few weeks ago and it flew completely under the radar.HRM, a new arch, reportedly has SOTA results on ARC-AGI 1 & 2 benchmarks with only 27 million parameters and ~1k training examples.— Alex MacCaw (@maccaw) July 25, 2025
ARC Prize on Twitter / X
OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design.… https://t.co/6Xk2Op0Sls— ARC Prize (@arcprize) July 30, 2026
Where’s my ten minute AGI?
Why don’t AIs automate more real-world tasks if they can handle 1-hour ones? Here are at least three fundamental reasons.

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.
François Chollet on Twitter / X
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3:1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents.2. Fine: general-purpose API settings that were not developed…— François Chollet (@fchollet) July 30, 2026
Letter to Arc members 2025
On Arc, its future, and the arrival of AI browsers — a moment to answer the largest questions you've asked us this past year.

Letter to Arc members 2025
On Arc, its future, and the arrival of AI browsers — a moment to answer the largest questions you've asked us this past year.

Arc Institute
Arc Institute is a independent nonprofit research organization headquartered in Palo Alto, California.

Arcee AI | Use MergeKit to Extract LoRA Adapters from any Fine-Tuned Model
We show you how to use Arcee's MergeKit to extract LoRA adapters from fine-tuned models, then leverage the Hugging Face Hub to create a library of general and task-specific LoRA adapters.

Import AI 447: The AGI economy; testing AIs with generated games; and agent ecologies
What might a superintelligence arcology be like?

Arcee AI | Announcing the Arcee Model Engine Public Beta
Get direct access to the small language models (SLMs) that power Arcee Orchestra, our new end-to-end, SLM-powered agentic AI platform. Sign up for the public beta of the Arcee Model Engine today.
.webp)
Arcee AI | March is Merge Madness
To celebrate Arcee’s recent merger with mergekit, we’re bringing you a month of resources and knowledge on model merging.
.webp)
The Receding User: Designing for Absence
AGI isn’t the ‘Holy Grail’ for women in AI. It’s gender-purpose AI and it’s already here - Fast Company
Women AI Leaders are shaping the future of the technology in ways that are more purpose-driven.
