







Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.
Open Models Inference for Coding · Umans AI
Hosted Kimi K3, GLM 5.2, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.

OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.

Hatice Ozen on Twitter / X
PSA: @OpenAI is putting the Open back in OpenAI and @GroqInc has Day 0 support. 🤗GPT-OSS 20B and 120B, hybrid-reasoning models with built-in browser search and code execution are now live for instant inference.P.S. We've also launched OpenAI Responses API compatibility. pic.twitter.com/CK7StvMSpr— Hatice Ozen (@ozenhati) August 5, 2025
ggml
AI inference at the edge. ggml has 22 repositories available. Follow their code on GitHub.
Open models by OpenAI
Advanced open-weight reasoning models to customize for any use case and run anywhere.

Dria on Twitter / X
Introducing Inference Arena v2.0.An agentic experience that searches, analyzes, and delivers insights about LLM inference.When we first launched, our goal was simple: make it easier for developers to compare models, engines, and hardware without digging through scattered… pic.twitter.com/fgWgos48lW— Dria (@driaforall) September 30, 2025
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

Run DeepSeek-R1 Dynamic 1.58-bit
DeepSeek R-1 is the most powerful open-source reasoning model that performs on par with OpenAI's o1 model. Run the 1.58-bit Dynamic GGUF version by Unsloth.

High Performance AI Lab
High Performance AI Lab builds open inference systems and publishes the conditions behind every number — device, model, quant, and rep count.

gpt-oss:20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show | TechCrunch
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

OpenAI's open source LLM is a reasoning model, coming Next Thursday!
1.1K votes, 257 comments. 756K subscribers in the LocalLLaMA community. Subreddit to discuss locally hostable AI.
Latency optimization | OpenAI API
Improve latency across a wide variety of LLM-related use cases.
