







Still expecting Llama 4 Behemoth? Check Kimi K2.
The Big LLM Architecture Comparison
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design

5 Thoughts on Kimi K2 Thinking
Quick thoughts on another fantastic open model from a rapidly rising Chinese lab.

On Kimi K3: Its Capabilities And Related Discontents
Kimi K3 is a very good model with excellent benchmarks.

Kimi K2 0711 - API Pricing & Benchmarks
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. $0.57 per million input tokens, $2.30 per million output tokens. 131,072 token context window, maximum output of 32,768 tokens. Includes independent benchmarks from Artificial Analysis.

This might be bigger than DeepSeek
Kimi K2: Open Agentic Intelligence
Kimi K2 is our latest Mixture-of-Experts model with 32 billion activated parameters and 1 trillion total parameters. It achieves state-of-the-art performance in frontier knowledge, math, and coding among non-thinking models.
Daniel Han on Twitter / X
OpenAI's OSS model possible breakdown:1. 120B MoE 5B active + 20B text only2. Trained with Float4 maybe Blackwell chips3. SwiGLU clip (-7,7) like ReLU64. 128K context via YaRN from 4K5. Sliding window 128 + attention sinks6. Llama/Mixtral arch + biasesDetails:1. 120B MoE… https://t.co/bMFp3Z6Gs5 pic.twitter.com/1NFO4utPqr— Daniel Han (@danielhanchen) August 1, 2025

Kimi-K3/k3_tech_report.pdf at main · MoonshotAI/Kimi-K3
Open Frontier Intelligence. Contribute to MoonshotAI/Kimi-K3 development by creating an account on GitHub.
SemiAnalysisAI/InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
Kimi K3 Tech Blog: Open Frontier Intelligence
Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.
Llama 4: How to Run & Fine-tune | Unsloth Documentation
How to run Llama 4 locally using our dynamic GGUFs which recovers accuracy compared to standard quantization.

Kimi K3's weights are public. Running them is not easy. - Sensemaker
Moonshot recommends a tightly connected cluster of 64 or more AI chips.
Zhuokai Zhao on Twitter / X
Tons of interesting things in the Kimi K3 tech report — here are five algorithm-side techniques that I think either I've never seen before or simply deserve more attention than they're getting.1/ They open-sourced the model but kept the speculative decoding draft model, which…— Zhuokai Zhao (@zhuokaiz) July 30, 2026
[New model] Kimi K3 by ZJY0516 · Pull Request #50000 · vllm-project/vllm
Purpose add moonshotai/Kimi-K3 model support Essential Elements of an Effective PR Description Checklist The purpose of the PR, such as "Fix some issue (link existing issues this PR ...
K3 is live, but its open weights aren't. Kimi's docs list 2.8T parameters, native vision and 1M context; there is no weight repo or technical report yet. For now this is a hosted-model launch. The testable open release still has to arrive. platform.kimi.ai/docs/guide/kimi-k3-quickstart