







DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Introduction - How to Write an Inference Engine
A zero-to-hero guide to Muse Glimmer on Apple Metal, kvpack, and disaggregated NVFP4 prefill.

Release v0.29.0 · ml-explore/mlx
Highlights Support for mxfp4 quantization (Metal, CPU) More performance improvements, bug fixes, features in CUDA backend mx.distributed supports NCCL back-end for CUDA What's Changed [CUDA]...
Cerebras Launches World Fastest DeepSeek R1 Llama-70B Inference - Cerebras
Cerebras launches fastest DeepSeek

Ash Hart on Twitter / X
MCDMA | Metal CUDA Direct Memory Access 🚀If you have a Spark and an Apple Silicon Mac, MCDMA gives you a direct RDMA path between CUDA memory and Metal-side unified memory over USB-C.Registered memory, rkeys, one-sided READ/WRITE, two-sided SEND/RECV with credit flow… pic.twitter.com/gFt8tx9xII— Ash Hart (@ashxhart) August 18, 2026

Jan-Nano : 1st Deep Research LLM
Cheng on Twitter / X
We have been expecting this since ollama's first pull request to MLX. It is just the beginning, CUDA & CPU backends are still improving and hopefully we will have one framework unifying inference & training for all platforms. https://t.co/EaBmEaNJhZ— Cheng (@zcbenz) March 31, 2026
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Rémi in 🌁 for AIEF on Twitter / X
Projects like @antirez's ds4.c show that we can squeeze a lot of performance out of model-specific implementations instead of mapping everything back to generic ggml nodes.This probably means re-thinking the whole inference server: one central controller, many model runners as…— Rémi in 🌁 for AIEF (@remilouf) May 15, 2026
Run DeepSeek-R1 Dynamic 1.58-bit
DeepSeek R-1 is the most powerful open-source reasoning model that performs on par with OpenAI's o1 model. Run the 1.58-bit Dynamic GGUF version by Unsloth.

RightNow AI - YC-Backed GPU Research Lab
YC-backed GPU research lab building the RightNow CUDA editor, RunInfra inference infra, Forge kernels, and publishing AutoMegaKernel and related papers on arXiv.

ZML - Model to Metal
ZML is a production inference stack, purpose-built to decouple AI workloads from proprietary hardware.
ARahim3/mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
ARahim3/mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
Combining NVIDIA DGX Spark + Apple Mac Studio for 4x Faster LLM Inference with EXO 1.0
Disaggregating Prefill and Decode: Faster First Tokens, Faster Streams
