







We're back to iterating on Ensu. The latest release brings Gemma 4, on-device voice transcription, and faster image queries.
Unsloth AI on Twitter / X
Run Gemma 3n locally with our Dynamic GGUFs!✨@Google's Gemma 3n supports audio, vision, video & text and the 4B model fits on 8GB RAM for fast local inference.Fine-tuning is also supported in Unsloth.Gemma-3n-E4B GGUF: https://t.co/PliynxoKQc https://t.co/wMFWLjNaDR pic.twitter.com/lxsMNDmkW8— Unsloth AI (@UnslothAI) June 26, 2025

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.

Simon Willison on Twitter / X
I'm really impressed by the new Gemma 3nI tried a 7.5GB model from Ollama and a 15GB model through mlx-vlm - they seem very capable, and this is the first model of that size I've tried that can handle both image AND audio input in addition to text! https://t.co/hiR3qGW387— Simon Willison (@simonw) June 26, 2025
Gemma 4 Technical Report
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

A Visual Guide to Gemma 4 12B
An in-depth explainer to Gemma 4 12B; a unified, encoder-free multimodal model!

INIYSA on Twitter / X
Apple's on-device SLM, while not as strong in simple multilingual tasks as Google's Gemma 3 4B (3.3GB, 5GB in memory), seems to outperform Qwen3 4B or Phi 4 mini reasoning. It's very very impressive, especially considering its extremely reduced size of around ~1.4GB— INIYSA (@lafaiel) June 10, 2025
Gemma 3n model overview | Google AI for Developers
Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. This model includes innovations in parameter-efficient processing, including Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture that provides the flexibility to reduce compute and memory requirements. These models feature audio input handling, as well as text and visual data.

Google for Developers Blog - News about Web, Mobile, AI and Cloud
Learn how to build with Gemma 3n, a mobile-first architecture, MatFormer technology, Per-Layer Embeddings, and new audio and vision encoders.

Matt Mireles on Twitter / X
Introducing...Gemma 4 Multimodal Fine-Tuner for Apple Silicon- LoRA fine-tunning toolkit for Gemma LLM- runs locally on macOS via PyTorch and Metal- streams data from Google Cloud to your machine- fine-tune on audio, image and text- easy-to-use CLI wizardIf you want… pic.twitter.com/UduROxoxPU— Matt Mireles (@mattmireles) April 7, 2026

Ivan Fioravanti ᯅ on Twitter / X
Playing with Qwen3-TTS is and MLX-Audio locally on Mac Studio M3 Ultra 🔥Amazing model by @Alibaba_Qwen and great work by @Prince_Canuma bringing this magic to MLX!Tuning the voice with instructions feels like magic!Command and prompts to run it are in the video. pic.twitter.com/qIfWKOedXr— Ivan Fioravanti ᯅ (@ivanfioravanti) January 23, 2026
Gemma 3n: How to Run & Fine-tune | Unsloth Documentation
Run Google's new Gemma 3n locally with Dynamic GGUFs on llama.cpp, Ollama, Open WebUI and fine-tune with Unsloth!

Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

gemma3n
Gemma 3n models are designed for efficient execution on everyday devices such as laptops, tablets or phones.

Running LLMs on your iPhone: 40 tok/s Gemma 4 with MLX — Adrien Grondin, Locally AI
Superwhisper — AI Voice to Text for macOS, Windows & iOS
AI powered voice to text for macOS, Windows, and iOS. Dictate in any app with offline and cloud speech recognition, 100+ languages, and custom AI modes.