







Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

gemma3n
Gemma 3n models are designed for efficient execution on everyday devices such as laptops, tablets or phones.

Gemma 3n: How to Run & Fine-tune | Unsloth Documentation
Run Google's new Gemma 3n locally with Dynamic GGUFs on llama.cpp, Ollama, Open WebUI and fine-tune with Unsloth!

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier
Capacity to train strong models is proliferating.

Open models in perpetual catch-up
The open-closed gap, distillation, innovation timescales, how open models win, specialized models, what’s missing, etc.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI

A 10 year old Xeon is all you need - point.free
Or running Gemma 4 on a 2016 Xeon with no GPU, 25 flags, 128 GB of DDR3, and a 25B-parameter MoE.

gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

Performance Explorer — oMLX
Compare model performance across context lengths with community benchmark data.

Open models are decelerationist - Erlend’s notes
people and planet need open models to win
Open models by OpenAI
Advanced open-weight reasoning models to customize for any use case and run anywhere.

Kasey Zhang on Twitter / X
We fine-tuned the model on 100K real-world code commits. On the test dataset (subset of CommitPackFT), foundation model performance ranged between 0.77-0.93 reward score. Osmosis-Apply-1.7B achieved a 0.98 reward score - while also being more affordable and ~10X faster!(Our… pic.twitter.com/CWS2APXoww— Kasey Zhang (@_WEEXIAO) July 3, 2025

gpt-oss:20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

Gemma 3n model overview | Google AI for Developers
Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. This model includes innovations in parameter-efficient processing, including Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture that provides the flexibility to reduce compute and memory requirements. These models feature audio input handling, as well as text and visual data.

Hint: it's not benchmark scores.
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/mtp