







Run Google's new Gemma 3n locally with Dynamic GGUFs on llama.cpp, Ollama, Open WebUI and fine-tune with Unsloth!
Unsloth AI on Twitter / X
Run Gemma 3n locally with our Dynamic GGUFs!✨@Google's Gemma 3n supports audio, vision, video & text and the 4B model fits on 8GB RAM for fast local inference.Fine-tuning is also supported in Unsloth.Gemma-3n-E4B GGUF: https://t.co/PliynxoKQc https://t.co/wMFWLjNaDR pic.twitter.com/lxsMNDmkW8— Unsloth AI (@UnslothAI) June 26, 2025

gemma3n
Gemma 3n models are designed for efficient execution on everyday devices such as laptops, tablets or phones.

Google for Developers Blog - News about Web, Mobile, AI and Cloud
Learn how to build with Gemma 3n, a mobile-first architecture, MatFormer technology, Per-Layer Embeddings, and new audio and vision encoders.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
Matt Mireles on Twitter / X
Introducing...Gemma 4 Multimodal Fine-Tuner for Apple Silicon- LoRA fine-tunning toolkit for Gemma LLM- runs locally on macOS via PyTorch and Metal- streams data from Google Cloud to your machine- fine-tune on audio, image and text- easy-to-use CLI wizardIf you want… pic.twitter.com/UduROxoxPU— Matt Mireles (@mattmireles) April 7, 2026

unslothai/unsloth
Unified web UI for training and running open models like Qwen, DeepSeek, gpt-oss and Gemma locally.
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Introducing Gemma 3n – the latest Google open model for accessible AI, featuring unique flexibility, privacy, and expanded multimodal capabilities on mobile devices.

Gemma 3n model overview | Google AI for Developers
Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. This model includes innovations in parameter-efficient processing, including Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture that provides the flexibility to reduce compute and memory requirements. These models feature audio input handling, as well as text and visual data.


A 10 year old Xeon is all you need - point.free
Or running Gemma 4 on a 2016 Xeon with no GPU, 25 flags, 128 GB of DDR3, and a 25B-parameter MoE.

Google for Developers Blog - News about Web, Mobile, AI and Cloud
Explore Gemma 3 270M, a compact, energy-efficient AI model for task-specific fine-tuning, offering strong instruction-following and production-ready quantization.

Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

Unsloth - Train and Run Models Locally
Unsloth is an open-source, no-code web UI for training, running and exporting open models in one unified local interface.

Tiles version 0.4.17 Alpha 21 has been released. The new default model is Gemma 4 12B, using Unsloth’s Q4_K_M GGUF. Added quantization tags to Modelfiles and automatic MTP decoding. Upgraded Pi to upstream v0.84.2, plus minor fixes for macOS notarization and GGUF model downloads. Release notes ↗
Shared chat session by @tiles.run | Tiles
chat.tiles.runGemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/mtp