







ByteShape's first public release of ShapeLearn-quantized models.
Reverse-engineering GGUF | Post-Training Quantization
ModelScope on Twitter / X
🤯 400 Token/S on a MacBook? Yes, you read that right!Shaohong Chen just fine-tuned the Qwen3-0.6B LLM in under 2 minutes using Apple's MLX framework. This is how you turn your MacBook into a serious LLM development rig. A step-by-step guide and performance metrics inside! 🧵… pic.twitter.com/31Cmycy8Mh— ModelScope (@ModelScope2022) October 13, 2025

Benjamin Marie on Twitter / X
Quantization and Qwen3-VL✔️ AutoRound W4A16 (INT4)✔️ AutoRound NVFP4 (llm compressor format)❌ support by vLLM (tried stable and dev releases)VLMs are now very easy to quantize. And then we can’t use the models with fast inference frameworks like vLLM and SGLang.It…— Benjamin Marie (@bnjmn_marie) November 5, 2025
Llama 4: How to Run & Fine-tune | Unsloth Documentation
How to run Llama 4 locally using our dynamic GGUFs which recovers accuracy compared to standard quantization.

Deep double descent: where bigger models and more data hurt*
Deep double descent: where bigger models and more data hurt* , Nakkiran, Preetum, Kaplun, Gal, Bansal, Yamini, Yang, Tristan, Barak, Boaz, Sutskever, Ilya
The Kaitchup Index: A Leaderboard for LLMs and Their Quantized Versions
Comparing formats like GGUF, GPTQ, and AWQ, with different bitwidths

A 30B Qwen Model Walks Into a Raspberry Pi… and Runs in Real Time
ByteShape's device-optimized release of Qwen3-30B-A3B-Instruct-2507.
A Visual Guide to Quantization
Exploring memory-efficient techniques for LLMs

Artur Chakhvadze on Twitter / X
We are releasing our first quantized checkpoints for the Qwen3.5 series of models, co-designed jointly with our inference engine to achieve maximum possible performance on Apple hardwareStarting from 0.8B, 2B and 4B modelshttps://t.co/2R8BdhAfzv— Artur Chakhvadze (@norpadon) June 8, 2026
8 Graphs Telling Today's Story of Open Models
Qwen on Twitter / X
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:- Autonomous coding: 10+ days of… pic.twitter.com/e3YFj2hqcT— Qwen (@Alibaba_Qwen) August 3, 2026

Overview - GroqDocs
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.

Alex Cheema on Twitter / X
.@karpathy shouted out my work on @exolabs at @ycombinator AI SUS!“we use LLMs similarly to mainframes in the ‘70s - compute is timeshared by having a slice in the batch dimension. models will compress over time, and with this we’ll be able to run more on-device” pic.twitter.com/UKSrquQGAL— Alex Cheema (@alexocheema) June 18, 2025

Qwen on Twitter / X
Qwen3.8 is launching and going open-weight soon!🌐With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.You don't have to wait to… pic.twitter.com/JS3ID73IYS— Qwen (@Alibaba_Qwen) July 19, 2026

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier
Capacity to train strong models is proliferating.
