







OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design.… https://t.co/6Xk2Op0Sls— ARC Prize (@arcprize) July 30, 2026
François Chollet on Twitter / X
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3:1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents.2. Fine: general-purpose API settings that were not developed…— François Chollet (@fchollet) July 30, 2026
Alex MacCaw on Twitter / X
I suspect generalized reasoning was solved just a few weeks ago and it flew completely under the radar.HRM, a new arch, reportedly has SOTA results on ARC-AGI 1 & 2 benchmarks with only 27 million parameters and ~1k training examples.— Alex MacCaw (@maccaw) July 25, 2025
François Chollet on Twitter / X
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.In fact,…— François Chollet (@fchollet) September 3, 2026


ChatGPT Voice can keep talking while it works - Sensemaker
OpenAI’s GPT-Live splits live conversation from slower search, reasoning, and agent work in the background.
OpenAI partners with Cerebras
OpenAI partners with Cerebras to add 750MW of high-speed AI compute, reducing inference latency and making ChatGPT faster for real-time AI workloads.

ARC-AGI-3
ARC-AGI-3 is the first interactive reasoning benchmark for AI agents—play as humans and build agents that learn in novel environments.

Letting an AI remember tripled its puzzle score - Sensemaker
OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.
How OpenAI delivers low-latency voice AI at scale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.

Latency optimization | OpenAI API
Improve latency across a wide variety of LLM-related use cases.

Diogo Almeida on Twitter / X
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev• 20-200x faster• 40-400x… pic.twitter.com/JSybNG2BKJ— Diogo Almeida (@CompleteSkeptic) September 15, 2026
I Benchmarked OpenAI Memory vs LangMem vs Letta (MemGPT) vs Mem0 for Long-Term Memory: Here’s How They Stacked Up
145 votes, 53 comments. Lately, I’ve been testing memory systems to handle long conversations in agent setups, optimizing for: Factual consistency…
What makes a great ChatGPT app | OpenAI Developers
How to build capabilities that make conversations better.

OpenAI’s Sam Altman on Building the ‘Core AI Subscription’ for Your Life