







Intelligence × decode speed × memory × quantization retention, under real iPhone limits. Same protocol for every model; Apple's built-in FM on the board.
Part 3: iPhone Hardware and How It Powers On-Device AI
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

Artificial Analysis on Twitter / X
We benchmarked Apple's new On-Device model: trails most Gemma and Qwen on-device suitable models but still very usefulGPQA Diamond performance trailed models that are suitable for on-device use such as the smaller Gemma models (3n E4B, 4B, 12B) and Qwen3 models (1.7B, 4B, 8B).… pic.twitter.com/wMrNM7yinL— Artificial Analysis (@ArtificialAnlys) June 20, 2025

Part 4: Brief history of Apple ML Stack
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.

The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
Introducing Apple’s On-Device and Server Foundation Models
At the 2024 Worldwide Developers Conference, we introduced Apple Intelligence, a personal intelligence system integrated deeply into iOS 18…

LLM Rankings | OpenRouter
LLM rankings and AI leaderboard based on benchmarks and real usage data from millions of users. See which AI models developers actually use.
The Kaitchup Index: A Leaderboard for LLMs and Their Quantized Versions
Comparing formats like GGUF, GPTQ, and AWQ, with different bitwidths

Introducing LFM2: The Fastest On-Device Foundation Models on the Market | Liquid AI
Today, we release LFM2, a new class of Liquid Foundation Models (LFMs) that sets a new standard in quality, speed, and memory efficiency for on-device deployment. Built on a hybrid architecture, LFM2 delivers 200% faster decode and prefill performance than Qwen3 and Gemma 3 on CPU. It also significantly outperforms models in each size class on instruction-following and function calling—the core capabilities that make LLMs reliable for building AI agents.

On-Device LLM Throughput Calculator - a Hugging Face Space by FL33TW00D-HF
This tool estimates and visualizes the throughput of Large Language Models on devices with memory bandwidth constraints. Users input device and model configurations, and the tool generates a plot s...
A Visual Guide to Quantization
Exploring memory-efficient techniques for LLMs

Atila on Twitter / X
Most advanced Apple Intelligence model device support: pic.twitter.com/mZ9rCVDijn— Atila (@atiorh) June 8, 2026

Every Apple M Chip Explained – M1, M2, M3, M4, M5
Melange | On-device AI for Mobile
Select. Benchmark. Deploy | End-to-end on-device AI deployment tool for mobile devs.
Apple’s biggest announcement today was Memory Integrity Enforcement · Victor Wynne
Today’s Apple event introduced Memory Integrity Enforcement on iPhone 17, the most significant memory safety upgrade in consumer OS history.
