







Introducing Muse Spark 1.3, with max reasoning for challenging reasoning and agentic tasks and improved real-world usability.
Muse Spark 1.3 (max) - Intelligence, Performance & Price Analysis | Artificial Analysis
Analysis of Meta's Muse Spark 1.3 (max) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Muse Spark 1.3 (xhigh) - Intelligence, Performance & Price Analysis | Artificial Analysis
Analysis of Meta's Muse Spark 1.3 (xhigh) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
A taxonomy for next-generation reasoning models
Where we've been and where we're going with RLVR.

Introduction - How to Write an Inference Engine
A zero-to-hero guide to Muse Glimmer on Apple Metal, kvpack, and disaggregated NVFP4 prefill.

gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

gpt-oss:20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

Muse Glimmer: Meta’s 30B Model Built for Efficient Inference
Inside Meta’s 30B local reasoning model and its tiny KV cache

Crafting a good (reasoning) model
A recent talk I gave on model training, reasoning, and the next frontier.

ARC-AGI-3
ARC-AGI-3 is the first interactive reasoning benchmark for AI agents—play as humans and build agents that learn in novel environments.

Prime Intellect - The Open Stack for Self-Improving Agents
The compute and infrastructure platform for you to train, evaluate, and deploy your own agentic models.

Prime Intellect - The Open Stack for Self-Improving Agents
The compute and infrastructure platform for you to train, evaluate, and deploy your own agentic models.

Gemma 4 Technical Report
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

Inkling: Our Open-Weights Model
Our first open-weights model: multimodal, Mixture-of-Experts, with controllable reasoning effort. Available to fine-tune on Tinker.

Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
Introducing Laguna S 2.1
Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.

Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl
Notes from my Thoughtworks colleagues on AI-assisted software delivery
