







GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.In fact,…— François Chollet (@fchollet) September 3, 2026
ARC-AGI-3
ARC-AGI-3 is the first interactive reasoning benchmark for AI agents—play as humans and build agents that learn in novel environments.

Alex MacCaw on Twitter / X
I suspect generalized reasoning was solved just a few weeks ago and it flew completely under the radar.HRM, a new arch, reportedly has SOTA results on ARC-AGI 1 & 2 benchmarks with only 27 million parameters and ~1k training examples.— Alex MacCaw (@maccaw) July 25, 2025
François Chollet on Twitter / X
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3:1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents.2. Fine: general-purpose API settings that were not developed…— François Chollet (@fchollet) July 30, 2026
GPT-6 Astra: A new generation of intelligence
Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.

gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

ARC Prize on Twitter / X
OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design.… https://t.co/6Xk2Op0Sls— ARC Prize (@arcprize) July 30, 2026
gpt-oss:20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.
OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems
Math is hard.

Prime Agent: A self-improving RLM agent
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, surpassing the reported human expert baseline.

What the hell happened with AGI timelines in 2026?
AI #181: Astra Goes Cyber Critical
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.

GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.

GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.

Model Release Notes | OpenAI Help Center
We’re beginning the rollout of GPT-5.6 Sol in ChatGPT, our flagship reasoning model for complex work across coding, research, science, cybersecurity, computer use, and design.GPT-5.6 Sol is rolling out to eligible paid ChatGPT plans. Free, Go, and logged-out users are not included. Availability may vary during rollout, and managed-workspace access can depend on administrator settings. Availability for other GPT-5.6 family models varies by product and plan; check the model picker or current rate card in the product you use.

Our team just shipped Fugu-Ultra v1.1! 🐡 By dynamically orchestrating the latest frontier models, we pushed performance up by 7.9 points. We are now beating Fable 5 in complex coding and reasoning tasks without even having Fable 5 in our agent pool. Collective intelligence is the future.
Sakana AI
Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work. Today, we’re releasing Fugu-Ultra v1.1 → sakana.ai/fugu Upgraded to incorporate the latest frontier models.