







Today, we're previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple's machine learning framework.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

Geek Lite on Twitter / X
在 Apple Silicon Mac 上本地运行 LLM 推理服务,提供比 Ollama 和 llama.cpp 更快的 OpenAI 兼容 API,同时原生支持工具调用和提示缓存。https://t.co/IXet9GW7x0Rapid-MLX 用 Apple 自家的 MLX 框架做推理,搭了个 FastAPI 服务跑 OpenAI 兼容 API。在 Apple Silicon 上比 Ollama 快 2-4 倍,靠… pic.twitter.com/d9QUbAWJSZ— Geek Lite (@QingQ77) May 3, 2026

Cheng on Twitter / X
We have been expecting this since ollama's first pull request to MLX. It is just the beginning, CUDA & CPU backends are still improving and hopefully we will have one framework unifying inference & training for all platforms. https://t.co/EaBmEaNJhZ— Cheng (@zcbenz) March 31, 2026
Exploring LLMs with MLX and the Neural Accelerators in the M5 GPU
Mac with Apple silicon is increasingly popular among AI developers and researchers interested in using their Mac to experiment with the…

Osaurus — Own Your AI on Apple Silicon
Own your AI: local-first agents with memory, tools, and identity on Apple Silicon. Offline, open source, and API-compatible with OpenAI, Anthropic, and Ollama.

Osaurus — Own Your AI on Apple Silicon
Own your AI: local-first agents with memory, tools, and identity on Apple Silicon. Offline, open source, and API-compatible with OpenAI, Anthropic, and Ollama.

clandestine.eth 🦇🔊 on Twitter / X
Heterogeneous acceleration on Apple Silicon achieved.ANE + GPU running in parallel.Mirror SD with DFlash, ported to MLX — targeting ANE + GPU simultaneously.The M-series was designed for this. We just hadn't unlocked it yet. pic.twitter.com/raSH0CMN4V— clandestine.eth 🦇🔊 (@0xClandestine) April 15, 2026

Part 4: Brief history of Apple ML Stack
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

mzau/broke-cluster
A Poor Man's Apple Silicon LLM Cluster — tuned for MLX, scalable without shame.
Run Ollama with NVIDIA GPU in Proxmox VMs and LXC containers
Learn how to run Ollama with an NVIDIA GPU in Proxmox for an enhanced AI experience in your home lab and great chat performance

Mirai for MacOS: A faster, simpler alternative to Ollama and LM Studio
Chat with your favorite AI models directly on your Mac. Privately and securely. Built natively for macOS and Apple Silicon.

omlx/docs/experimental/dflash_mlx_integration.md at main · jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar - jundot/omlx

Apple stumbled into succes with MLX
201 votes, 76 comments. Qwen3-next 80b-a3b is out in mlx on hugging face, MLX already supports it. Open source contributors got this done within 2…