







People who think current AI use is unsustainable often rely on the claim that inference GPUs only last “three years at the most” under load1. The idea here is that once the AI bubble money drains away, current infrastructure will rapidly become obsolete, and there won’t be enough money floating around to buy a whole slate of brand-new GPUs. Inference costs would thus rapidly become way too expensive for current AI products to make any financial sense.
GPU Pricing — Live Platform Rates | Vast.ai
Live GPU pricing on Vast.ai. Prices set by supply and demand across 40+ data centers. On-demand, interruptible, or reserved — find the right GPU at the right price.
Open Source AI Inference Benchmark | InferenceX
Compare AI inference performance across GPUs and frameworks. Real benchmarks on NVIDIA GB200, B200, AMD MI355X, and more. Free, open-source, continuously updated.
Rohan Paul on Twitter / X
Google is trying to win AI by making compute cheap, not by beating Nvidia on raw speed.Nvidia sells GPUs to clouds with a big 70%+ margin that sits on top of manufacturing and R&D cost and raises cloud prices.Google builds TPUs for itself at near manufacturing cost, adds no… https://t.co/aSgWRf0HY7 pic.twitter.com/T3Fzc6czwg— Rohan Paul (@rohanpaul_ai) November 25, 2025

Where AI Startups Scale to Production
Discover the most efficient way to build, tune and run your AI models and applications on top-notch NVIDIA® GPUs.

Wafer - Ship the fastest inference in the world
Autonomous AI agents that profile, diagnose, and optimize GPU inference across your entire stack — from kernels to models to production pipelines.

The Other Bubble
Buried in the 8000 words I wrote last week was a worrying story — that Microsoft considered drastic measures to free up capacity in its US-based servers for GPUs to power the AI boom. In an email shared with me by a source from earlier this year, Microsoft's senior leadership team

Import AI 464: Fables writes GPU kernels; AI automation; and analog computation
Is this the beginning of a new world?

NVIDIA Rubin CPX Accelerates Inference Performance and Efficiency for 1M+ Token Context Workloads | NVIDIA Technical Blog
Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent memory, and long-horizon context—enabling…

AI Doesn't Have ROI
If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year, or $7 a month, and in return you get a weekly newsletter that’s usually anywhere from 5,000 to 18,000 words, including vast, detailed analyses of NVIDIA, Anthropic and OpenAI’s

The AI Industry Is Losing
If you liked this piece, you should subscribe to my premium newsletter. It’s $70 a year, or $7 a month, and in return you get a weekly newsletter that’s usually anywhere from 5,000 to 18,000 words, including vast, detailed analyses of NVIDIA, Anthropic and OpenAI’s

Three-Dimensional Unit Economics: The Compute Cost of Retention
Why AI-native apps need a new financial framework. And a new metric to run it.

Advancing AI Infrastructure for Agentic AI with NVIDIA DOCA In-Silicon Security | NVIDIA Technical Blog
The AI era is driving a new class of infrastructure: AI factories that transform data into intelligence for autonomous AI agents operating at unprecedented scale. Powered by accelerated computing…

Spending on AI Is at Epic Levels. Will It Ever Pay Off?
Tech companies are pouring hundreds of billions into data centers, taking on heavy debt, but current revenue is relatively tiny. Critics warn of a new dot-com bubble.
We're Not Building AI Features for the Money
From the Zed Blog: Why Zed invests in AI, and the future we're building toward.
Confidential AI Cloud · Private Inference on GPU TEE | Phala
Phala Cloud delivers confidential AI on TEE-protected GPUs — private LLM inference, sealed agents, and verifiable compute on Intel TDX + NVIDIA H100/H200.
