







With a smaller 256-chip footprint per Pod, TPU v5e is optimized to be a high-value product for transformer, text-to-image, and Convolutional Neural Network (CNN) training, fine-tuning, and serving. For more information about using Cloud TPU v5e for serving, see Inference using v5e.
TPU architecture | Google Cloud Documentation
Tensor Processing Units (TPUs) are application specific integrated circuits (ASICs) designed by Google to accelerate machine learning workloads. Cloud TPU is a Google Cloud service that makes TPUs available as a scalable resource.
Run a calculation on a Cloud TPU VM using PyTorch | Google Cloud Documentation
Learn how to create a Cloud TPU, install PyTorch and run a simple calculation on a Cloud TPU.
We're launching two specialized TPUs for the agentic era.
The eighth generation of Google’s TPU includes two specialized chips that will power the future of AI.

Expanding our use of Google Cloud TPUs and Services
Announcing a dramatic increase in Anthropic's compute resources
Cloud TPU quotas | Google Cloud Documentation
This document lists the quotas that apply to Cloud TPU. For information about Cloud TPU pricing, see Cloud TPU pricing.
Google AI Plans with Cloud Storage - Google One
Explore Google AI Plans. Access our most advanced AI, generate videos from text, and secure cloud storage.
Rohan Paul on Twitter / X
Google is trying to win AI by making compute cheap, not by beating Nvidia on raw speed.Nvidia sells GPUs to clouds with a big 70%+ margin that sits on top of manufacturing and R&D cost and raises cloud prices.Google builds TPUs for itself at near manufacturing cost, adds no… https://t.co/aSgWRf0HY7 pic.twitter.com/T3Fzc6czwg— Rohan Paul (@rohanpaul_ai) November 25, 2025

Pricing | Runpod
GPU cloud computing at up to 80% less than hyperscalers. Explore pricing for on-demand Pods, Serverless, Clusters, and Network Storage.

Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

google-ai-edge/LiteRT
LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
Google for Developers Blog - News about Web, Mobile, AI and Cloud
LiteRT is the universal framework for on-device AI. The production stack delivers 1.4x faster cross-platform GPU performance, streamlined NPU acceleration, and superior GenAI support for open models like Gemma.

Getting started with Transformers and TPU using PyTorch
Learn how to get started with Hugging Face Transformers and TPUs using PyTorch, fine-tune a BERT model for Text Classification using the newest Google Cloud TPUs.
Google Cloud release notes | Google Cloud Documentation
The following release notes cover the most recent changes over the last 60 days. For a comprehensive list of product-specific release notes, see the individual product release note pages.
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Explore Gemma 3 270M, a compact, energy-efficient AI model for task-specific fine-tuning, offering strong instruction-following and production-ready quantization.
