François Chollet on Twitter / X
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3:1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents.2. Fine: general-purpose API settings that were not developed…— François Chollet (@fchollet) July 30, 2026
ARC Prize on Twitter / X
OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design.… https://t.co/6Xk2Op0Sls— ARC Prize (@arcprize) July 30, 2026
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.

Letting an AI remember tripled its puzzle score - Sensemaker
OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.