







Today, we release LFM2, a new class of Liquid Foundation Models (LFMs) that sets a new standard in quality, speed, and memory efficiency for on-device deployment. Built on a hybrid architecture, LFM2 delivers 200% faster decode and prefill performance than Qwen3 and Gemma 3 on CPU. It also significantly outperforms models in each size class on instruction-following and function calling—the core capabilities that make LLMs reliable for building AI agents.
Liquid AI — Device-native foundation models.
Liquid AI is an efficiency-first foundation model company. We build highly capable, compute-optimized models that bring intelligence to any device and medium of choice.

Liquid AI on Twitter / X
Today, we release LFM2.5-350M. Agentic loops at 350M parameters.A 350M model trained for reliable data extraction and tool use, where models at this scale typically struggle.<500MB when quantized, built for environments where compute, memory, and latency are constrained.🧵 pic.twitter.com/zZPKzcCwH9— Liquid AI (@liquidai) March 31, 2026

Liquid AI Launches LEAP and Liquid Apollo: The Easiest Way to Build with On-Device AI | Liquid AI
Today marks a pivotal milestone in the evolution of edge AI. Liquid AI is thrilled to announce LEAP v0, our first developer-ready platform for on-device AI deployment—and Liquid Apollo, a lightweight iOS-native application built to showcase and stress-test small foundation models directly on your phone.

Liquid <> .txt Collaboration | Liquid AI
Faster and more accurate function calling on the edge with .txt’s structured outputs and Liquid Foundation Models

Google for Developers Blog - News about Web, Mobile, AI and Cloud
LiteRT is the universal framework for on-device AI. The production stack delivers 1.4x faster cross-platform GPU performance, streamlined NPU acceleration, and superior GenAI support for open models like Gemma.

google-ai-edge/LiteRT
LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks
Which AI model writes the best code? We rank every major LLM — open and closed source — across SWE-bench, HumanEval, LiveCodeBench, and Terminal-Bench coding benchmarks. Compare the best LLMs for coding, software engineering, and programming.

The Ma of a New Machine – Scott Jenson
The current Silicon Valley flex is trading notes on your favorite new AI tools over lunch. Each week brings another one to explore. Some of these tools are very impressive; I’ve been able to reply to someone with an alternative UI design in less than a minute, and they were dumbfounded: “How did you do that so fast?”

Part 4: Brief history of Apple ML Stack
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

LLMs Can Now Write GPU Kernels That Beat torch.compile - Break AI Scaling Limits in 7 Days
We're now seeing multi-agent systems that take your PyTorch code and produce CUDA or Triton kernels with 2x to 14x speedups over torch.compile(mode='max-autotune-no-cudagraphs'). Not on toy benchmarks. On real models like Llama-3.1-8B, Whisper, and Stable Diffusion. Learn proven techniques to shift the scaling law intercept and achieve 10-50% performance gains.

Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AI
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
David Hendrickson on Twitter / X
🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched. Holo 3.1 is H Company’s (🇫🇷) new local computer-use agent model that beats Qwen3.5-397B, Kimi-K2.5, and Sonnet 4.6!Since it is built for local deployment → ⬩ Runs fully on your machine… https://t.co/CpOEsuWN2k pic.twitter.com/w39iOh7cO1— David Hendrickson (@TeksEdge) June 2, 2026
