







OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design.… https://t.co/6Xk2Op0Sls— ARC Prize (@arcprize) July 30, 2026
François Chollet on Twitter / X
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3:1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents.2. Fine: general-purpose API settings that were not developed…— François Chollet (@fchollet) July 30, 2026
Alex MacCaw on Twitter / X
I suspect generalized reasoning was solved just a few weeks ago and it flew completely under the radar.HRM, a new arch, reportedly has SOTA results on ARC-AGI 1 & 2 benchmarks with only 27 million parameters and ~1k training examples.— Alex MacCaw (@maccaw) July 25, 2025
François Chollet on Twitter / X
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.In fact,…— François Chollet (@fchollet) September 3, 2026


ChatGPT Voice can keep talking while it works - Sensemaker
OpenAI’s GPT-Live splits live conversation from slower search, reasoning, and agent work in the background.
OpenAI partners with Cerebras
OpenAI partners with Cerebras to add 750MW of high-speed AI compute, reducing inference latency and making ChatGPT faster for real-time AI workloads.

ARC-AGI-3
ARC-AGI-3 is the first interactive reasoning benchmark for AI agents—play as humans and build agents that learn in novel environments.

Letting an AI remember tripled its puzzle score - Sensemaker
OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.
How OpenAI delivers low-latency voice AI at scale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.

Latency optimization | OpenAI API
Improve latency across a wide variety of LLM-related use cases.

I Benchmarked OpenAI Memory vs LangMem vs Letta (MemGPT) vs Mem0 for Long-Term Memory: Here’s How They Stacked Up
145 votes, 53 comments. Lately, I’ve been testing memory systems to handle long conversations in agent setups, optimizing for: Factual consistency…
What makes a great ChatGPT app | OpenAI Developers
How to build capabilities that make conversations better.

OpenAI’s Sam Altman on Building the ‘Core AI Subscription’ for Your Life
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.

nWave – The AI-augmented Framework for Software Crafters
nWave, an open AI framework that replaces ad‑hoc prompting with repeatable SDLC waves: human intent + agent execution, so teams ship faster, safer, less rework.
