







Not all software should change at the same speed. This has always been true, but it's easy to forget when tools make change frictionless. Generative AI dram…
The Universal Execution Layer for AI
Optimize any AI model on any engine, across all hardware. Dria’s topology-aware compiler and peer-to-peer runtime merge CPUs, GPUs, NPUs & chiplets into one fabric—maximising utilisation, cutting inference cost and ending vendor lock-in.

The Ma of a New Machine – Scott Jenson
The current Silicon Valley flex is trading notes on your favorite new AI tools over lunch. Each week brings another one to explore. Some of these tools are very impressive; I’ve been able to reply to someone with an alternative UI design in less than a minute, and they were dumbfounded: “How did you do that so fast?”

Regenerative Software - The Phoenix Architecture
The Growth OS Map: Building Defensible Loops in the AI Era
The 8 loops and the 5-layer architecture needed to replace fragile funnels with a resilient GTM engine.

Heaps do lie: debugging a memory leak in vLLM. | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
Scott Jenson – Exploring the world beyond mobile
Fast AI requires slow thinking The current Silicon Valley flex is trading notes on your favorite new AI tools over lunch. Each week brings another one to explore. Some of these tools are very impressive; I’ve been able to reply to someone with an alternative UI design in less than a minute, and they were […]
The Generative Stack - The Phoenix Architecture
Trying to find the best tool or platform for generative software in 2026 is a mistake that could haunt you for decades

Code Was Never the Asset - The Phoenix Architecture
Why AI makes the hidden economics of software unavoidable
The Shape of Memory Benchmarks
Why the familiar memory benchmarks are outdated, how the agent-native work looks today and why design your own.

LukeW | Common AI Product Issues
At this point, almost every software domain has launched or explored AI features. Despite the wide range of use cases, most of these implementations have been t...

Teaching AI to Optimize AI Models for Edge Deployment
How our agent, Möbius automated a Core ML port in ~12 h (vs. 2 weeks), hit 0.99998 parity, and made it 3.5× faster, while staying on the CPU.

raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Alex Komoroske on Twitter / X
The real AI battle isn't about who has the best model—it's about who controls your context. I just published a piece on why keeping these layers separate is the most important design decision of the AI era: https://t.co/6bljxg7KIt— Alex Komoroske (@komorama) June 16, 2025
AI & Alignment Raw coding speed isn't the bottleneck. Alignment is the bottleneck. That seems to be a zeitgeist-y theme lately. If you're using AI to code, maybe you're feeling it. You can code more and faster. And clearly a boatload of other developers are doing that too. But software doesn't…
AI & Alignment
chriscoyier.net