







The first megakernel for hybrid DeltaNet/Attention LLMs. 413 tok/s at 1.87 tok/J on a 2020 RTX 3090, matching M5 Max efficiency at 1.8x throughput.
apple-silicon-llm-bench/results/complete_results.html at main · AlexHiesch/apple-silicon-llm-bench
Systematic LLM inference benchmark for Apple Silicon: 8 backends, 7 models, 791 measurements - AlexHiesch/apple-silicon-llm-bench
EXO Labs reveals that they have been working with Apple for the past year on low-latency RDMA networking over TB5 which allows a cluster of 4 x M5 Ultra Mac Studios to scale to an aggregate memory bandwidth of 4.8TB/s
the tiny corp on Twitter / X
Qwen 3.5 27B getting 18.5 tok/s on Mac Mini with external 7900XTX. It should be able to be 3x faster than this with work, SSM stuff is still in PR. Hopefully Mac eGPU support brings in devs. pic.twitter.com/2aMkUpXY1S— the tiny corp (@__tinygrad__) April 1, 2026

New Apple Silicon Drive record... For LLMs
Benchmarks | EXO
Transparent benchmarks for LLMs tested on real hardware. Coming soon.
mzau/broke-cluster
A Poor Man's Apple Silicon LLM Cluster — tuned for MLX, scalable without shame.
Optimizing On-Device Inference for Apple Silicon
A custom local engine that improves prefill and decode throughput

Alex Cheema on Twitter / X
4 x M5 Max MacBooks clustered with RDMA:512GB @ 2456GB/s, $20k, 560W, quiet.Find me a better deal, that I can buy today. https://t.co/pWZsHnv0Kl— Alex Cheema (@alexocheema) May 13, 2026
The Big LLM Architecture Comparison
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design

0xSero on Twitter / X
I told y’all this is the move. Heterogenous hardware is the way forward. Large cheap pools of mixed memory + specialized accelerators (Nvidia GPUs, DGX Spark, Cerebras wafers) The next year will be dominated by solutions that split the stack. - 3000$ for a used Mac Studio… https://t.co/zMlScSnJ0X— 0xSero (@0xSero) April 1, 2026
MTPLX
2.24x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
On-Device LLM Throughput Calculator - a Hugging Face Space by FL33TW00D-HF
This tool estimates and visualizes the throughput of Large Language Models on devices with memory bandwidth constraints. Users input device and model configurations, and the tool generates a plot s...
Doriandarko/MLX-GRPO
A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.
GMKtec EVO-X2 AI Mini PC AMD Ryzen™ AI Max+ 395
EVO-X2 AMD Ryzen™ Al Max+ 395 Mini PC, the World’s First Windows 11 AI+ PC APU Supporting 70B LLM, Experience relentless power with the AMD Ryzen™ AI Max+ 395. With 16 cores and 32 threads built on the groundbreaking Zen 5 architecture, it clocks up to 5.1GHz. Crafted using TSMC’s advanced 4nm FinFET process and armed with a massive 16MB L2 and 64MB L3 cache, it delivers seamless multitasking and unrivaled parallel processing.

AMD and Anthropic Announce Strategic Partnership to Deploy Up to 2 Gigawatts of AMD Instinct MI450 Series GPUs
News Highlights Anthropic to deploy up to 2 gigawatts of MI450 Series GPUs in AMD Helios rack-scale solutions, with deployment of the first…...
