







Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.
Ollama's new engine for multimodal models· Ollama Blog
Ollama now supports new multimodal models with its new engine.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
qwen3.5:27b
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.

Simon Willison on Twitter / X
I'm really impressed by the new Gemma 3nI tried a 7.5GB model from Ollama and a 15GB model through mlx-vlm - they seem very capable, and this is the first model of that size I've tried that can handle both image AND audio input in addition to text! https://t.co/hiR3qGW387— Simon Willison (@simonw) June 26, 2025
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.

Gemma 4 Technical Report
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework. As these models are rapidly evolving toward general-purpose instruction following across diverse and complex tasks, a key frontier is evaluating their crosslingual and multimodal capabilities over both short- and long-form inputs. However, existing benchmarks fall short in evaluating these dimensions jointly: they are often limited to English, mostly focus on a single modality at a time, rely on short-form inputs, or lack human annotations--hindering comprehensive assessment of model performance across languages, modalities, and task complexity. To address these gaps, we introduce MCIF (Multimodal Crosslingual Instruction Following), the first crosslingual human-annotated benchmark based on scientific talks on NLP and beyond. MCIF evaluates instruction following in crosslingual, multimodal settings over different input lengths and spans four macro-tasks: recognition, translation, question answering, and summarization. It covers three core modalities (speech, vision, and text) and four diverse languages (English, German, Italian, and Chinese), fully aligned across all dimensions. This parallel design enables a systematic evaluation of MLLMs' abilities to interpret instructions across languages and effectively integrate multimodal contextual information. Our benchmarking and analysis of 23 models highlight universal challenges across modalities and tasks, indicating substantial room for improvement in future MLLMs development. MCIF is released under CC-BY 4.0 license to promote open research.

Understanding Multimodal LLMs
An introduction to the main techniques and latest models

OhAPI | One API. Infinite Adult AI Experiences.
One unified multimodal stack for text, image, voice, video and digital twins. Built for scale and monetisation.

🚨 We're very happy to introduce TRIBE v2: a foundation model of the human brain's responses to sight, sound, and language. Leveraging 1,000+ hours of fMRI across 720 subjects, it generalizes… | Stéphane d'Ascoli | 16 comments
🚨 We're very happy to introduce TRIBE v2: a foundation model of the human brain's responses to sight, sound, and language. Leveraging 1,000+ hours of fMRI across 720 subjects, it generalizes zero-shot to new stimuli, tasks and people, finetunes efficiently, and enables in-silico experiments. ❓How does it work? Stemming from our v1, which won the Algonauts 2025 challenge, TRIBE v2 combines video, audio, and language embeddings to predict brain activity for any brain, then adapts to each individual. Key results: 📊 High-quality predictions — TRIBE v2 predicts brain activity across cortical and subcortical regions, significantly better than standard linear models, with a log-linear scaling law and no plateau in sight. 🎯 Zero-shot generalization — Without retraining, the predictions of TRIBE v2 are more correlated with group-averaged brain responses than almost any individual fMRI scan! A short finetuning step vastly improves over linear models trained, from scratch, on each individual. 🧪 In-silico experiments — Can we do useful experiments with TRIBE v2? Yes: classic vision and language paradigms replicate in-silico. It zero-shot recovers the FFA, PPA, EBA, VWFA, Broca's lateralization, and syntactic responses in STG — all without training on these artificial tasks. 🔍 Interpretability & multimodality — ICA on the weights rediscovers known functional networks (auditory, language, motion, default mode, visual) from naturalistic data alone. Ablating modalities further maps how vision, audition, and language integrate, with the largest gains at the temporo-parietal-occipital junction. 🧠 This effort is a step toward a foundation model of the human brain. Much remains to be understood, but we hope this opens a path for neuroscience, AI, and medical research alike. All code, weights, and a live demo are open — find it useful or mistaken in some conditions? Let us know, new test cases can only help improving this effort. 📄 Paper: https://lnkd.in/e7cbunJp 💻 Code: https://lnkd.in/ebwBVuJp ▶️ Demo: https://lnkd.in/eEUVxP4S 🤗 Model: https://lnkd.in/e2T8nPJP Joint work with Jérémy RAPIN, Yohann Benchetrit, Teon Brooks, Katie Begany, Joséphine Raugel, Hubert Banville and Jean-Rémi King. 🙏 Special thanks to Elisa Cascardi, Diego Marcos, Dominic Giardini, AI at Meta, and the open-source and neuroscience communities (in particular Lune Bellec and Bertrand Thirion for the amazing Courtois NeuroMod and IBC datasets) | 16 comments on LinkedIn
👋 Jan on Twitter / X
Introducing Jan-v2-VL, a multimodal agent built for long-horizon tasks.Jan-v2-VL executes 49 steps without failure, while the base model stops at 5 and other similar-scale VLMs stop between 1 and 2.It achieves longer, stable task execution in your browser without accuracy… pic.twitter.com/rwzOIGhlFq— 👋 Jan (@jandotai) November 13, 2025
README.md · MiniMaxAI/MiniMax-H3 at 73372e6cf53e414edd3ab03e357717fb0602e758
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Interaction Models: A Scalable Approach to Human-AI Collaboration
Interaction models move beyond turn-based AI interfaces by handling multimodal, real-time collaboration natively across audio, video, and text.
Kimi K3 Tech Blog: Open Frontier Intelligence
Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.
Why MLX — Prince Canuma, Neywa Labs