







We are excited to share a preview of our distributed inference stack — engineered for consumer GPUs and the 100ms latencies of the public internet—plus a research roadmap that scales it into a planetary-scale inference engine.
Dria on Twitter / X
Introducing Inference Arena v2.0.An agentic experience that searches, analyzes, and delivers insights about LLM inference.When we first launched, our goal was simple: make it easier for developers to compare models, engines, and hardware without digging through scattered… pic.twitter.com/fgWgos48lW— Dria (@driaforall) September 30, 2025
Overview - GroqDocs
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.

The Inference Shift
Agentic inference is going to be different than the inference we use today, and it will change compute infrastructure because speed won’t matter when humans aren’t involved.

Hatice Ozen on Twitter / X
PSA: @OpenAI is putting the Open back in OpenAI and @GroqInc has Day 0 support. 🤗GPT-OSS 20B and 120B, hybrid-reasoning models with built-in browser search and code execution are now live for instant inference.P.S. We've also launched OpenAI Responses API compatibility. pic.twitter.com/CK7StvMSpr— Hatice Ozen (@ozenhati) August 5, 2025
Open-Source Agentic Inference Benchmark | InferenceX
Compare AgentX, InferenceX's long-context, multi-turn coding scenario, with fixed-sequence AI inference across chips and frameworks. Public NVIDIA and AMD runs update when configurations change.

Can we billionaire-proof inference? - Graze Newsletter
A small, federated compute co-op called co/core — and the wider experiment it's a part of.
Demystifying llm-d and vLLM: The race to production
Learn how vLLM and llm-d work together for efficient and scalable large language model (LLM) inference. Discover the benefits of disaggregated scaling, expert-parallel scheduling, and KV cache-aware routing.

Cheng on Twitter / X
We have been expecting this since ollama's first pull request to MLX. It is just the beginning, CUDA & CPU backends are still improving and hopefully we will have one framework unifying inference & training for all platforms. https://t.co/EaBmEaNJhZ— Cheng (@zcbenz) March 31, 2026
Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

Public AI Inference Utility
A nonprofit, open-source service to make public and sovereign AI models more accessible.
Open Models Inference for Coding · Umans AI
Hosted Kimi K3, GLM 5.2, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.

SambaNova | The Fastest AI Inference Platform
Discover SambaNova - the complete AI platform delivering the fastest AI inference, fine-tuning, and scalable solutions for agentic AI easily integrated into existing data center infrastructures.

Muse Glimmer: Meta’s 30B Model Built for Efficient Inference
Inside Meta’s 30B local reasoning model and its tiny KV cache

OpenAI's open source LLM is a reasoning model, coming Next Thursday!
1.1K votes, 257 comments. 756K subscribers in the LocalLLaMA community. Subreddit to discuss locally hostable AI.
Fully Sharded Data Parallel
We’re on a journey to advance and democratize artificial intelligence through open source and open science.