







Free · 32 concurrent · hosted by selimaktas on LocalMaxxing.
Qwen on Twitter / X
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:- Autonomous coding: 10+ days of… pic.twitter.com/e3YFj2hqcT— Qwen (@Alibaba_Qwen) August 3, 2026

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Friday’s big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba’s Qwen research lab. I’ve been looking forward to this one: 27B is an …

A 30B Qwen Model Walks Into a Raspberry Pi… and Runs in Real Time
ByteShape's device-optimized release of Qwen3-30B-A3B-Instruct-2507.
qwen3.5:27b
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.

To The Qwen Team, Kindly Contribute to Qwen3-Next GGUF Support!
442 votes, 126 comments. If you haven't noticed already, Qwen3-Next hasn't yet been supported in llama.cpp, and that's because it comes with a custom…
Qwen on Twitter / X
Qwen3-TTS is officially live. We’ve open-sourced the full family—VoiceDesign, CustomVoice, and Base—bringing high quality to the open community.- 5 models (0.6B & 1.8B)- Free-form voice design & cloning- Support for 10 languages- SOTA 12Hz tokenizer for high compression-… pic.twitter.com/BSWpaYoZWj— Qwen (@Alibaba_Qwen) January 22, 2026

Qwen
Qwen is a family of large language models developed by Alibaba Cloud. Many Qwen models are distributed under the free and open-source Apache 2.0 license, the source-available Qwen License, or the non-commercial Qwen Research License; other proprietary Qwen models are served through Alibaba Cloud.

Alibaba Unveils Qwen3.8-Max: Its Largest and Most Capable Flagship Model to Date - Alibaba Cloud
Alibaba unveils Qwen3.8-Max, its most powerful model with 2.4 trillion parameters, excelling in coding, research, and visual intelligence.

Release Auto compaction (preview) + LAN Remote Access · unslothai/unsloth
Thanks for the support for Qwen3.8-27B and Unsloth Desktop last week! For this release, we merged 200+ PRs to introduce many new features, fixes including: Auto Compaction (Experimental) for longe...
Localmaxxing
About half of agent tasks can run on a local 35B model. The real advantage isn't cost or privacy — it's latency. 2.1x faster means more iteration cycles per session.
Awni Hannun on Twitter / X
The latest mlx-lm is out and it has continuous batching with mlx_lm.server! Added by @angeloskath Check-out the video of 4 simultaneous requests running with Qwen3 30B on the same M2 Ultra: https://t.co/o9sFC3k4DN— Awni Hannun (@awnihannun) December 3, 2025
Artur Chakhvadze on Twitter / X
We are releasing our first quantized checkpoints for the Qwen3.5 series of models, co-designed jointly with our inference engine to achieve maximum possible performance on Apple hardwareStarting from 0.8B, 2B and 4B modelshttps://t.co/2R8BdhAfzv— Artur Chakhvadze (@norpadon) June 8, 2026
Paul Frazee — Solving for scale in open networks
atomic.chat on Twitter / X
Google Turbo Quant running Locally in Atomic ChatMacBook Air M4 16 GBModel: QWEN3.5-9BContext window: 50000Summarising 20000 words in just seconds..You can do 3x larger context window, processing 3x faster than before! pic.twitter.com/FRYkXCGjQb— atomic.chat (@atomic_chat_hq) March 27, 2026
Georgi Gerganov on Twitter / X
I think the consensus is that Qwen3.5 is a step change so atm I would recommend explore that, given that it covers a range of sizes suitable for all devices.Note that the main issues that people currently unknowingly face with local models mostly revolve around the harness and…— Georgi Gerganov (@ggerganov) March 30, 2026