







The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark. We will bill based on the total number of input and output tokens by the model.
Simple Pricing | Machine Learning Infrastructure | Deep Infra
We provide only pay-what-you-use pricing with no long-term contracts or upfront costs for our machine learning models and infrastructure. Learn more!

Kimi K2 0711 - API Pricing & Benchmarks
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. $0.57 per million input tokens, $2.30 per million output tokens. 131,072 token context window, maximum output of 32,768 tokens. Includes independent benchmarks from Artificial Analysis.
tokens are getting more expensive
"language models will get cheaper by 10x" will not save ai subscriptions from the short squeeze

A broken pricing paradigm
A token is not a fixed unit of cost Variance in usage creates an interconnected pricing and scaling issue Anjali Shrivastava anjali.shrivastava99@gmail.com | @anjali_shriva THIS ESSAY IS NOW LIVE! A token is not a fixed unit of cost PART I: High variance in AI demand breaks unit economics and re...



Groq On-demand Pricing for Tokens-as-a-Service
Groq powers leading openly-available AI models. View the pricing of our core models including GPT-OSS, Kimi K2, Qwen3 32B, and more.

AI Inference Pricing, EU Hosted, Per Token | TensorX
Transparent, pay-as-you-go pricing for private AI inference on TensorX. No lock-in, EU-hosted, with zero data retention and an OpenAI-compatible API.


BytePlus Free Trial: 4K AI Image/Video Models + LLM Tokens | 200 Free Images
BytePlus free trial, Seedream 4.0, Seedance 1.0, DeepSeek V3.1, GPT-OSS-120B, free 4K AI images, AI video generation, BytePlus LLM tokens

Flagship Model Kimi K3 Pricing - Kimi API Platform
Kimi K3 is our flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window and industry-leading intelligence. The Kimi API Platform provides K3, K2.7 Code, K2.6 and other large language model APIs, supporting long context, multimodal understanding, and Tool Calling.

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.

The Kaitchup – AI on a Budget | Benjamin Marie | Substack
Weekly tutorials and news on adapting large language models (LLMs) to your tasks and hardware using the most recent techniques and models. The Kaitchup proposes a collection of 180+ AI notebooks regularly updated. Click to read The Kaitchup – AI on a Budget, by Benjamin Marie, a Substack publication with tens of thousands of subscribers.

Open Models Inference for Coding · Umans AI
Hosted Kimi K3, GLM 5.2, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.

SemiAnalysis on Twitter / X
Recently, we purchased one of each Anthropic/OpenAI subscription plan and randomly ran long horizon coding tasks until we exhausted the weekly limit. It's widely believed that a $200/month plan maxes out at ~$2000/month worth of tokens (assuming API pricing). However, we found… pic.twitter.com/1e0zFhbFuo— SemiAnalysis (@SemiAnalysis_) June 10, 2026
