







An in-depth explainer to Gemma 4 12B; a unified, encoder-free multimodal model!
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.

Gemma 4 Technical Report
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
Gemma 3n model overview | Google AI for Developers
Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. This model includes innovations in parameter-efficient processing, including Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture that provides the flexibility to reduce compute and memory requirements. These models feature audio input handling, as well as text and visual data.

Gemma 4: Byte for byte, the most capable open models
Gemma 4: our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.

Matt Mireles on Twitter / X
Introducing...Gemma 4 Multimodal Fine-Tuner for Apple Silicon- LoRA fine-tunning toolkit for Gemma LLM- runs locally on macOS via PyTorch and Metal- streams data from Google Cloud to your machine- fine-tune on audio, image and text- easy-to-use CLI wizardIf you want… pic.twitter.com/UduROxoxPU— Matt Mireles (@mattmireles) April 7, 2026

Understanding Multimodal LLMs
An introduction to the main techniques and latest models

Ensu continues
We're back to iterating on Ensu. The latest release brings Gemma 4, on-device voice transcription, and faster image queries.

Unsloth AI on Twitter / X
Run Gemma 3n locally with our Dynamic GGUFs!✨@Google's Gemma 3n supports audio, vision, video & text and the 4B model fits on 8GB RAM for fast local inference.Fine-tuning is also supported in Unsloth.Gemma-3n-E4B GGUF: https://t.co/PliynxoKQc https://t.co/wMFWLjNaDR pic.twitter.com/lxsMNDmkW8— Unsloth AI (@UnslothAI) June 26, 2025

qwen3.5:27b
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.

Ollama's new engine for multimodal models· Ollama Blog
Ollama now supports new multimodal models with its new engine.

Simon Willison on Twitter / X
I'm really impressed by the new Gemma 3nI tried a 7.5GB model from Ollama and a 15GB model through mlx-vlm - they seem very capable, and this is the first model of that size I've tried that can handle both image AND audio input in addition to text! https://t.co/hiR3qGW387— Simon Willison (@simonw) June 26, 2025
Inside the World's Smartest Robot Brain [VLA]
Transformer Explainer: LLM Transformer Model Visually Explained
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework. As these models are rapidly evolving toward general-purpose instruction following across diverse and complex tasks, a key frontier is evaluating their crosslingual and multimodal capabilities over both short- and long-form inputs. However, existing benchmarks fall short in evaluating these dimensions jointly: they are often limited to English, mostly focus on a single modality at a time, rely on short-form inputs, or lack human annotations--hindering comprehensive assessment of model performance across languages, modalities, and task complexity. To address these gaps, we introduce MCIF (Multimodal Crosslingual Instruction Following), the first crosslingual human-annotated benchmark based on scientific talks on NLP and beyond. MCIF evaluates instruction following in crosslingual, multimodal settings over different input lengths and spans four macro-tasks: recognition, translation, question answering, and summarization. It covers three core modalities (speech, vision, and text) and four diverse languages (English, German, Italian, and Chinese), fully aligned across all dimensions. This parallel design enables a systematic evaluation of MLLMs' abilities to interpret instructions across languages and effectively integrate multimodal contextual information. Our benchmarking and analysis of 23 models highlight universal challenges across modalities and tasks, indicating substantial room for improvement in future MLLMs development. MCIF is released under CC-BY 4.0 license to promote open research.

Mckay Wrigley on Twitter / X
factorio is unironically the perfect tutorial for agentic coding systems btw. https://t.co/CCkveZ4Yzr— Mckay Wrigley (@mckaywrigley) December 27, 2025