







Gemma 3n models are designed for efficient execution on everyday devices such as laptops, tablets or phones.
Gemma 3n model overview | Google AI for Developers
Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. This model includes innovations in parameter-efficient processing, including Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture that provides the flexibility to reduce compute and memory requirements. These models feature audio input handling, as well as text and visual data.

Gemma 3n: How to Run & Fine-tune | Unsloth Documentation
Run Google's new Gemma 3n locally with Dynamic GGUFs on llama.cpp, Ollama, Open WebUI and fine-tune with Unsloth!

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Learn how to build with Gemma 3n, a mobile-first architecture, MatFormer technology, Per-Layer Embeddings, and new audio and vision encoders.

Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

A 10 year old Xeon is all you need - point.free
Or running Gemma 4 on a 2016 Xeon with no GPU, 25 flags, 128 GB of DDR3, and a 25B-parameter MoE.

Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency
We’re releasing Gemma 4 quantization-aware training checkpoints, reducing memory requirements and improving on-device performance.

Google for Developers Blog - News about Web, Mobile, AI and Cloud
Introducing Gemma 3n – the latest Google open model for accessible AI, featuring unique flexibility, privacy, and expanded multimodal capabilities on mobile devices.

Introducing Gemma 4 12B: a unified, encoder-free multimodal model
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.

Google for Developers Blog - News about Web, Mobile, AI and Cloud
Explore Gemma 3 270M, a compact, energy-efficient AI model for task-specific fine-tuning, offering strong instruction-following and production-ready quantization.


Unsloth AI on Twitter / X
Run Gemma 3n locally with our Dynamic GGUFs!✨@Google's Gemma 3n supports audio, vision, video & text and the 4B model fits on 8GB RAM for fast local inference.Fine-tuning is also supported in Unsloth.Gemma-3n-E4B GGUF: https://t.co/PliynxoKQc https://t.co/wMFWLjNaDR pic.twitter.com/lxsMNDmkW8— Unsloth AI (@UnslothAI) June 26, 2025

Matt Mireles on Twitter / X
Introducing...Gemma 4 Multimodal Fine-Tuner for Apple Silicon- LoRA fine-tunning toolkit for Gemma LLM- runs locally on macOS via PyTorch and Metal- streams data from Google Cloud to your machine- fine-tune on audio, image and text- easy-to-use CLI wizardIf you want… pic.twitter.com/UduROxoxPU— Matt Mireles (@mattmireles) April 7, 2026

Artificial Analysis on Twitter / X
We benchmarked Apple's new On-Device model: trails most Gemma and Qwen on-device suitable models but still very usefulGPQA Diamond performance trailed models that are suitable for on-device use such as the smaller Gemma models (3n E4B, 4B, 12B) and Qwen3 models (1.7B, 4B, 8B).… pic.twitter.com/wMrNM7yinL— Artificial Analysis (@ArtificialAnlys) June 20, 2025

Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/mtp