







Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Awni Hannun on Twitter / X
It's very cool that Apple shipped a 20B parameter on-device. You can't put 20B parameters in RAM at any reasonable precision. To make it work they are using pretty exotic architecture by today's standards.A small model predicts from the query (or prompt) which experts to load… pic.twitter.com/Zhe5HcbGuL— Awni Hannun (@awnihannun) June 9, 2026

Running local models on an M4 with 24GB memory | jola.dev
Experiments with getting usable outputs out of local models on a standard Macbook

Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency
We’re releasing Gemma 4 quantization-aware training checkpoints, reducing memory requirements and improving on-device performance.

Optimizing On-Device Inference for Apple Silicon
A custom local engine that improves prefill and decode throughput

apple-silicon-llm-bench/results/complete_results.html at main · AlexHiesch/apple-silicon-llm-bench
Systematic LLM inference benchmark for Apple Silicon: 8 backends, 7 models, 791 measurements - AlexHiesch/apple-silicon-llm-bench
Alex Cheema on Twitter / X
It’s kind of crazy but the shitstorm of supply chain issues has created a new best-in-class local AI deployment: M5 Max MacBook clusters.- The memory unit economics are great - each MacBook has 128GB @ 614GB/s for $5k- M5 Max added tensor cores (Apple Neural Accelerators) with… https://t.co/f8STQ0tLZs pic.twitter.com/FLm3oOnyEl— Alex Cheema (@alexocheema) May 14, 2026

Max Weinbach on Twitter / X
AFM Core Advanced on-device model running on A19 Pro is a sparse model. It's 20B parameters. It's fully Apple designed. It is an MoE but when it processes the prompt, it only loads the parameters needed and locks them in.If it's 20B parameters total, but on a specific…— Max Weinbach (@mweinbach) June 8, 2026
Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

Alex Cheema on Twitter / X
4 x M5 Max MacBooks clustered with RDMA:512GB @ 2456GB/s, $20k, 560W, quiet.Find me a better deal, that I can buy today. https://t.co/pWZsHnv0Kl— Alex Cheema (@alexocheema) May 13, 2026
Steeve Morin on Twitter / X
For instance, we are now loading from SSD at about 14-16GB/s *in a composable manner*.That means about 1.1s to load an 8B BF16 model. https://t.co/i0c7xu8HhP— Steeve Morin (@steeve) November 1, 2025
DHH on Twitter / X
Going to double down on the MacBook mission with Omarchy. We almost have perfect coverage for the vintage Intel era going from 2009-2020. There's a straight shot to get the M1 and M2 machines going too, even if it's a lot more work. But we'll do the work. We'll fix everything.— DHH (@dhh) August 22, 2026
Artur Chakhvadze on Twitter / X
We are releasing our first quantized checkpoints for the Qwen3.5 series of models, co-designed jointly with our inference engine to achieve maximum possible performance on Apple hardwareStarting from 0.8B, 2B and 4B modelshttps://t.co/2R8BdhAfzv— Artur Chakhvadze (@norpadon) June 8, 2026
A 10 year old Xeon is all you need - point.free
Or running Gemma 4 on a 2016 Xeon with no GPU, 25 flags, 128 GB of DDR3, and a 25B-parameter MoE.

jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
so that new Mac Studio and M5 Ultra have me thinking about 2030 now the wildcard on the $1200ish 2030 Mac Mini is RAM. who knows what the future holds for RAM but the trajectory is clear: a reasonably priced home computer is going to have absolutely sicko inference power within a few years now…
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/mtp