







A TurboQuant inference server

Overview - GroqDocs
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.

ggml
AI inference at the edge. ggml has 22 repositories available. Follow their code on GitHub.
Planetary-Scale Inference: Building a Distributed Inference Engine for the Public Internet
We are excited to share a preview of our distributed inference stack — engineered for consumer GPUs and the 100ms latencies of the public internet—plus a research roadmap that scales it into a planetary-scale inference engine.

Ironwood: The first Google TPU for the age of inference
We’re introducing Ironwood, our seventh-generation Tensor Processing Unit (TPU) designed to power the age of generative AI inference.

Datacurve | The data engine for frontier AI
Custom data for long-horizon reasoning, software engineering, and data science.

SambaNova | The Fastest AI Inference Platform
Discover SambaNova - the complete AI platform delivering the fastest AI inference, fine-tuning, and scalable solutions for agentic AI easily integrated into existing data center infrastructures.

Introduction - How to Write an Inference Engine
A zero-to-hero guide to Muse Glimmer on Apple Metal, kvpack, and disaggregated NVFP4 prefill.

TurboQuant: Redefining AI efficiency with extreme compression
Amir Zandieh, Research Scientist, and Vahab Mirrokni, VP and Google Fellow, Google Research

Benjamin Marie on Twitter / X
Quantization and Qwen3-VL✔️ AutoRound W4A16 (INT4)✔️ AutoRound NVFP4 (llm compressor format)❌ support by vLLM (tried stable and dev releases)VLMs are now very easy to quantize. And then we can’t use the models with fast inference frameworks like vLLM and SGLang.It…— Benjamin Marie (@bnjmn_marie) November 5, 2025
Open-Source Agentic Inference Benchmark | InferenceX
Compare AgentX, InferenceX's long-context, multi-turn coding scenario, with fixed-sequence AI inference across chips and frameworks. Public NVIDIA and AMD runs update when configurations change.