Letting an AI remember tripled its puzzle score - Sensemaker
OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.
ARC-AGI-3
ARC-AGI-3 is the first interactive reasoning benchmark for AI agents—play as humans and build agents that learn in novel environments.

Who earned the score? - Sensemaker
This week's AI claims blurred models, systems, simulations and people. The evidence becomes clearer when the tested subject comes first.
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.

AA-Omniscience: Knowledge and Hallucination Benchmark | Artificial Analysis
Compare AI model performance on AA-Omniscience: Knowledge and Hallucination Benchmark. A benchmark measuring factual recall and hallucination across various economically relevant domains.
Glimmer · Reproducible AI science
Glimmer turns a research project into a navigable knowledge graph you can explore, run, verify, and extend — reproducibly.
Why AI Coding Agents Forget — And How ArcticMem Fixes It
Explore ArcticMem, Snowflake’s persistent semantic memory system for AI coding agents. See how dual-tier memory improves benchmark pass rates to 73%.

Accelerating GPT-5.6 Sol Ultrafast with OpenAI
Cerebras powers OpenAI’s GPT-5.6 Sol Ultrafast in the OpenAI API, delivering frontier intelligence at real-time speeds for critical AI work.

claude-obsidian
Self-organizing AI second brain for Obsidian + Claude Code. Drop any source and Claude reads, links, and files it into one connected knowledge graph of plain Markdown you own. AI note-taking, personal knowledge management (PKM), and an open-source Notion alternative. Based on Karpathy's LLM Wiki pattern.
Alex MacCaw on Twitter / X
I suspect generalized reasoning was solved just a few weeks ago and it flew completely under the radar.HRM, a new arch, reportedly has SOTA results on ARC-AGI 1 & 2 benchmarks with only 27 million parameters and ~1k training examples.— Alex MacCaw (@maccaw) July 25, 2025
milla-jovovich/mempalace
The highest-scoring AI memory system ever benchmarked. And it's free.