







Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/mtp
Jun 11, 2026 at 4:28 PM
Unsloth AI on Twitter / X
Run Gemma 3n locally with our Dynamic GGUFs!✨@Google's Gemma 3n supports audio, vision, video & text and the 4B model fits on 8GB RAM for fast local inference.Fine-tuning is also supported in Unsloth.Gemma-3n-E4B GGUF: https://t.co/PliynxoKQc https://t.co/wMFWLjNaDR pic.twitter.com/lxsMNDmkW8— Unsloth AI (@UnslothAI) June 26, 2025

Gemma 3n: How to Run & Fine-tune | Unsloth Documentation
Run Google's new Gemma 3n locally with Dynamic GGUFs on llama.cpp, Ollama, Open WebUI and fine-tune with Unsloth!

A 10 year old Xeon is all you need - point.free
Or running Gemma 4 on a 2016 Xeon with no GPU, 25 flags, 128 GB of DDR3, and a 25B-parameter MoE.


gemma3n
Gemma 3n models are designed for efficient execution on everyday devices such as laptops, tablets or phones.

Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency
We’re releasing Gemma 4 quantization-aware training checkpoints, reducing memory requirements and improving on-device performance.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

MTPLX
2.24x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
INIYSA on Twitter / X
Apple's on-device SLM, while not as strong in simple multilingual tasks as Google's Gemma 3 4B (3.3GB, 5GB in memory), seems to outperform Qwen3 4B or Phi 4 mini reasoning. It's very very impressive, especially considering its extremely reduced size of around ~1.4GB— INIYSA (@lafaiel) June 10, 2025
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Explore Gemma 3 270M, a compact, energy-efficient AI model for task-specific fine-tuning, offering strong instruction-following and production-ready quantization.

Tiles version 0.4.17 Alpha 21 has been released. The new default model is Gemma 4 12B, using Unsloth’s Q4_K_M GGUF. Added quantization tags to Modelfiles and automatic MTP decoding. Upgraded Pi to upstream v0.84.2, plus minor fixes for macOS notarization and GGUF model downloads. Release notes ↗
Shared chat session by @tiles.run | Tiles
chat.tiles.runTiles version 0.4.19 Alpha 23 has been released. MTP is now opt-in, with a new --mtp flag for tiles run and persistent configuration under [llama]. Inference server warnings are now surfaced in the CLI, and fixed an issue where tiles update could install multiple versions. Release notes ↗
Shared chat session by @tiles.run | Tiles
chat.tiles.run