







LLMs and VLMs with MLX Swift
Exploring LLMs with MLX and the Neural Accelerators in the M5 GPU
Mac with Apple silicon is increasingly popular among AI developers and researchers interested in using their Mac to experiment with the…

Run LLMs locally on your Mac · mlx-optiq
Quantize, fine-tune and serve LLMs locally on Apple Silicon. MLX-native, no PyTorch, no cloud. On PyPI.

videlalvaro/ane-book
Production LLM inference on the Apple Neural Engine — a practitioner's guide, complete with converters, Swift runtimes, and validated model manifests
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
ethicalabs-ai/BlossomTuneLLM-MLX
Federated Fine-Tuning of LLMs on Apple Silicon with Flower.ai and MLX-LM
ModelScope on Twitter / X
🤯 400 Token/S on a MacBook? Yes, you read that right!Shaohong Chen just fine-tuned the Qwen3-0.6B LLM in under 2 minutes using Apple's MLX framework. This is how you turn your MacBook into a serious LLM development rig. A step-by-step guide and performance metrics inside! 🧵… pic.twitter.com/31Cmycy8Mh— ModelScope (@ModelScope2022) October 13, 2025

Doriandarko/MLX-GRPO
A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.
mlx-examples/stable_diffusion at main · ml-explore/mlx-examples
Examples in the MLX framework. Contribute to ml-explore/mlx-examples development by creating an account on GitHub.
The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)
Cross-Model Evaluation: kaish collection syntax across 7 LLMs (DeepSeek, Gemini, Claude, Gemma, GLM, Qwen)
Cross-Model Evaluation: kaish collection syntax across 7 LLMs (DeepSeek, Gemini, Claude, Gemma, GLM, Qwen) · GitHub

Release v0.29.0 · ml-explore/mlx
Highlights Support for mxfp4 quantization (Metal, CPU) More performance improvements, bug fixes, features in CUDA backend mx.distributed supports NCCL back-end for CUDA What's Changed [CUDA]...
mzau/broke-cluster
A Poor Man's Apple Silicon LLM Cluster — tuned for MLX, scalable without shame.
Awni Hannun on Twitter / X
The latest mlx-lm is out and it has continuous batching with mlx_lm.server! Added by @angeloskath Check-out the video of 4 simultaneous requests running with Qwen3 30B on the same M2 Ultra: https://t.co/o9sFC3k4DN— Awni Hannun (@awnihannun) December 3, 2025