







I recently spoke at AI Tinkerers Raleigh about hyperfast inference systems and how they’re fundamentally changing AI application design. If you haven’t heard of Cerebras (or however they pronounce it), you’re in for a treat—this is one of the most exciting areas of research in AI right now.

Scott Jenson – Exploring the world beyond mobile
Fast AI requires slow thinking The current Silicon Valley flex is trading notes on your favorite new AI tools over lunch. Each week brings another one to explore. Some of these tools are very impressive; I’ve been able to reply to someone with an alternative UI design in less than a minute, and they were […]
The Ma of a New Machine – Scott Jenson
The current Silicon Valley flex is trading notes on your favorite new AI tools over lunch. Each week brings another one to explore. Some of these tools are very impressive; I’ve been able to reply to someone with an alternative UI design in less than a minute, and they were dumbfounded: “How did you do that so fast?”

Cerebras
Cerebras is the go-to platform for fast and effortless AI training. Learn more at cerebras.ai.

Cerebras
Cerebras is the go-to platform for fast and effortless AI training. Learn more at cerebras.ai.

Cerebras
Cerebras is the go-to platform for fast and effortless AI training. Learn more at cerebras.ai.

Cerebras
Cerebras is the go-to platform for fast and effortless AI training. Learn more at cerebras.ai.

Accelerating GPT-5.6 Sol Ultrafast with OpenAI
Cerebras powers OpenAI’s GPT-5.6 Sol Ultrafast in the OpenAI API, delivering frontier intelligence at real-time speeds for critical AI work.

Ultra-Processed Information: AI and the Coming Deluge of Noise | Frankly 128
What We Learned from Letting AI Posttrain AI
We built a posttraining task that runs for 20 hours with the Tinker API. The core bottleneck is research intuition.

NVIDIA Rubin CPX Accelerates Inference Performance and Efficiency for 1M+ Token Context Workloads | NVIDIA Technical Blog
Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent memory, and long-horizon context—enabling…

Wafer - Ship the fastest inference in the world
Autonomous AI agents that profile, diagnose, and optimize GPU inference across your entire stack — from kernels to models to production pipelines.

TurboQuant: Redefining AI efficiency with extreme compression
Amir Zandieh, Research Scientist, and Vahab Mirrokni, VP and Google Fellow, Google Research

The Kaitchup – AI on a Budget | Benjamin Marie | Substack
Weekly tutorials and news on adapting large language models (LLMs) to your tasks and hardware using the most recent techniques and models. The Kaitchup proposes a collection of 180+ AI notebooks regularly updated. Click to read The Kaitchup – AI on a Budget, by Benjamin Marie, a Substack publication with tens of thousands of subscribers.

How Notion Cut Millions from Their Vector DB Bill Without Sacrificing Search Quality
something that has come up fairly recently with LLMs - for coding, specifically - is that it’s become a lot easier to burn stupefying amounts of tokens on stuff very fast, with agents running 24/7 or managing more agents (see: Yegge’s Gas Town) even with low inference costs that adds up in a hurry
Jesse Felder
‘While some cling to the promise of an AI “revolution,” the cost of adoption is proving a stubborn bottleneck. These developments also suggest that the economics of replacing human labor with AI may be more complicated than some early forecasts originally implied.’ fortune.com/2026/05/22/microsoft-ai-cost-…