







Alex Cheema on Twitter / X
4 x M5 Max MacBooks clustered with RDMA:512GB @ 2456GB/s, $20k, 560W, quiet.Find me a better deal, that I can buy today. https://t.co/pWZsHnv0Kl— Alex Cheema (@alexocheema) May 13, 2026
Alex Cheema on Twitter / X
It’s kind of crazy but the shitstorm of supply chain issues has created a new best-in-class local AI deployment: M5 Max MacBook clusters.- The memory unit economics are great - each MacBook has 128GB @ 614GB/s for $5k- M5 Max added tensor cores (Apple Neural Accelerators) with… https://t.co/f8STQ0tLZs pic.twitter.com/FLm3oOnyEl— Alex Cheema (@alexocheema) May 14, 2026

the tiny corp on Twitter / X
Qwen 3.5 27B getting 18.5 tok/s on Mac Mini with external 7900XTX. It should be able to be 3x faster than this with work, SSM stuff is still in PR. Hopefully Mac eGPU support brings in devs. pic.twitter.com/2aMkUpXY1S— the tiny corp (@__tinygrad__) April 1, 2026

Megakernel: Matching Apple Silicon Efficiency at 2x the Throughput on a RTX 3090
The first megakernel for hybrid DeltaNet/Attention LLMs. 413 tok/s at 1.87 tok/J on a 2020 RTX 3090, matching M5 Max efficiency at 1.8x throughput.

Ash Hart on Twitter / X
This is why local models must win. Cloud models refused to reverse-engineer Apple's RDMA protocol. DeepSeek and Nemotron said, "hold my beer". TBF, Nemotron needed KV Cache injection to comply lol. pic.twitter.com/z3l2azui7e— Ash Hart (@ashxhart) August 14, 2026
The Potential of M6 and M5 Ultra for Local AI on macOS
Earlier today, Apple unveiled the new generation of Mac mini and Mac Studio, featuring the latest entries in the Apple silicon family of chips: the M6, available in the Mac mini, and the M5 Ultra, exclusive to the Mac Studio. You can read more details about the announcement and related specs in John’s overview. As

New Apple Silicon Drive record... For LLMs
Is the Day of the Data Center About to Be Over?
Marco Arment's Setup as the Canary in the Coal Mine—or, Rather, as the 50 Mac Mini Server Farm Vastly More Efficient than the NVIDIA-Powered Cloud-Bound Hyperscalers...

Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute
Apple debuted M6 in the new Mac mini and M5 Ultra in the new Mac Studio, providing an extraordinary leap in performance and AI capabilities.

0xSero on Twitter / X
I told y’all this is the move. Heterogenous hardware is the way forward. Large cheap pools of mixed memory + specialized accelerators (Nvidia GPUs, DGX Spark, Cerebras wafers) The next year will be dominated by solutions that split the stack. - 3000$ for a used Mac Studio… https://t.co/zMlScSnJ0X— 0xSero (@0xSero) April 1, 2026
Awni Hannun on Twitter / X
It's very cool that Apple shipped a 20B parameter on-device. You can't put 20B parameters in RAM at any reasonable precision. To make it work they are using pretty exotic architecture by today's standards.A small model predicts from the query (or prompt) which experts to load… pic.twitter.com/Zhe5HcbGuL— Awni Hannun (@awnihannun) June 9, 2026

Running local models on an M4 with 24GB memory | jola.dev
Experiments with getting usable outputs out of local models on a standard Macbook

Mac Studio
The ultimate pro desktop. Powered by M4 Max and M3 Ultra for all-out performance and extensive connectivity. Built for Apple Intelligence.

First look at the DGX Spark
A local supercomputer between the size of a Mac mini and a Mac mini.

Combining NVIDIA DGX Spark + Apple Mac Studio for 4x Faster LLM Inference with EXO 1.0
Disaggregating Prefill and Decode: Faster First Tokens, Faster Streams

Added a new device to my @tiles.run cluster. Welcome to the lineup, M5 Pro with the 32 GB/1 TB spec.