







Docs for GGUF quantization (unofficial)
Thireus/GGUF-Tool-Suite
Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain will create a GGUF recipe tuned to your hardware within seconds — flexible model sizing and lowest achievable perplexity/kld for GGUF enthusiasts seeking precise and automated dynamic quant production.
Reverse-engineering GGUF | Post-Training Quantization
Benjamin Marie on Twitter / X
Quantization and Qwen3-VL✔️ AutoRound W4A16 (INT4)✔️ AutoRound NVFP4 (llm compressor format)❌ support by vLLM (tried stable and dev releases)VLMs are now very easy to quantize. And then we can’t use the models with fast inference frameworks like vLLM and SGLang.It…— Benjamin Marie (@bnjmn_marie) November 5, 2025
Quantization from the ground up | ngrok blog
A complete guide to what quantization is, how it works, and how it's used to compress large language models

Quantizer support in Mediabunny v1.52.0 | Mediabunny
Mediabunny v1.52.0 adds quantizer-based video encoding for AVC, HEVC, VP9, and AV1, enabling constant-quality video encoding.

Georgi Gerganov on Twitter / X
HuggingFace just shipped in-browser GGUF editingIt allows you to edit GGUF metadata in the comfort of your browser, without having to even download the full model. This feature is enabled via the Xet technology that makes partial file updates possible. pic.twitter.com/uKqsvfLHcT— Georgi Gerganov (@ggerganov) October 7, 2025
Llama 4: How to Run & Fine-tune | Unsloth Documentation
How to run Llama 4 locally using our dynamic GGUFs which recovers accuracy compared to standard quantization.

Run DeepSeek-R1 Dynamic 1.58-bit
DeepSeek R-1 is the most powerful open-source reasoning model that performs on par with OpenAI's o1 model. Run the 1.58-bit Dynamic GGUF version by Unsloth.

ggml
AI inference at the edge. ggml has 22 repositories available. Follow their code on GitHub.
The Kaitchup Index: A Leaderboard for LLMs and Their Quantized Versions
Comparing formats like GGUF, GPTQ, and AWQ, with different bitwidths

A 4-Bit Model and a 1-Bit Index
Running NVFP4 Nemotron on a CPU, then mapping every embedding-compression method at matched byte budgets. The two quantizations compose.

Cody - Sourcegraph docs
Documentation for Sourcegraph, the code intelligence platform.
Information: the Measure of All Things? Part I: Communication, Code and Computation - 3 Quarks Daily
by Yohan J. John
RightNow AI - YC-Backed GPU Research Lab
YC-backed GPU research lab building the RightNow CUDA editor, RunInfra inference infra, Forge kernels, and publishing AutoMegaKernel and related papers on arXiv.

Karpathy's LLM Wiki: The Complete Guide to His Idea File
Karpathy's follow-up gist went viral. Complete breakdown with implementation examples.
