







201 votes, 76 comments. Qwen3-next 80b-a3b is out in mlx on hugging face, MLX already supports it. Open source contributors got this done within 2…
Zhijian Liu on Twitter / X
🔥 DFlash x MLX is happening!Shoutout to @aryagm01 for the early work on this. We're building on the momentum. Native MLX support, more models (Qwen3.5), up to 4x faster. Lossless!👉 https://t.co/9CtLKDptNI pic.twitter.com/pMdaCH4oKi— Zhijian Liu (@zhijianliu_) April 15, 2026
Ivan Fioravanti ᯅ on Twitter / X
Playing with Qwen3-TTS is and MLX-Audio locally on Mac Studio M3 Ultra 🔥Amazing model by @Alibaba_Qwen and great work by @Prince_Canuma bringing this magic to MLX!Tuning the voice with instructions feels like magic!Command and prompts to run it are in the video. pic.twitter.com/qIfWKOedXr— Ivan Fioravanti ᯅ (@ivanfioravanti) January 23, 2026
Osaurus — All your AI. One app.
Chat with GPT-5, Claude, Llama, and more — or download local MLX models from Hugging Face. Supercharge Cursor with powerful tools. Open source and built for Mac.
Casper Hansen on Twitter / X
Qwen3.5 Small models about to release!Qwen3.5 9B, 4B, 2B, 0.8B, or something in between is possible.- imagine 9B beating Qwen3-Next-80B- or 4B beating Qwen3-VL-30B in multimodal reasoningBuying a GPU is starting to have high return of intelligence on investment— Casper Hansen (@casper_hansen_) March 1, 2026
Awni Hannun on Twitter / X
The latest mlx-lm is out and it has continuous batching with mlx_lm.server! Added by @angeloskath Check-out the video of 4 simultaneous requests running with Qwen3 30B on the same M2 Ultra: https://t.co/o9sFC3k4DN— Awni Hannun (@awnihannun) December 3, 2025
mzau/mlx-knife
ollama like cli tool for MLX models on huggingface (pull, rm, list, show, serve etc.)
Ollama is now powered by MLX on Apple Silicon in preview· Ollama Blog
Today, we're previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple's machine learning framework.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
Artur Chakhvadze on Twitter / X
We are releasing our first quantized checkpoints for the Qwen3.5 series of models, co-designed jointly with our inference engine to achieve maximum possible performance on Apple hardwareStarting from 0.8B, 2B and 4B modelshttps://t.co/2R8BdhAfzv— Artur Chakhvadze (@norpadon) June 8, 2026
Qwen (Qwen)
Org profile for Qwen on Hugging Face, the AI community building the future.
qwen3.5:27b
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.

Qwen on Twitter / X
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:- Autonomous coding: 10+ days of… pic.twitter.com/e3YFj2hqcT— Qwen (@Alibaba_Qwen) August 3, 2026

ModelScope on Twitter / X
🤯 400 Token/S on a MacBook? Yes, you read that right!Shaohong Chen just fine-tuned the Qwen3-0.6B LLM in under 2 minutes using Apple's MLX framework. This is how you turn your MacBook into a serious LLM development rig. A step-by-step guide and performance metrics inside! 🧵… pic.twitter.com/31Cmycy8Mh— ModelScope (@ModelScope2022) October 13, 2025

Qwen3-Coder: Agentic Coding in the World
GITHUB HUGGING FACE MODELSCOPE DISCORD Today, we’re announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we’re excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct — a 480B-parameter Mixture-of-Experts model with 35B active parameters which supports the context length of 256K tokens natively and 1M tokens with extrapolation methods, offering exceptional performance in both coding and agentic tasks. Qwen3-Coder-480B-A35B-Instruct sets new state-of-the-art results among open models on Agentic Coding, Agentic Browser-Use, and Agentic Tool-Use, comparable to Claude Sonnet 4.
mzau/broke-cluster
A Poor Man's Apple Silicon LLM Cluster — tuned for MLX, scalable without shame.