







Meet Pipette, an open-source platform for reproducible on-device AI benchmarks across models, quantization, runtimes and hardware.
Melange | On-device AI for Mobile
Select. Benchmark. Deploy | End-to-end on-device AI deployment tool for mobile devs.
Part 3: iPhone Hardware and How It Powers On-Device AI
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

Part 4: Brief history of Apple ML Stack
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

Alex Cheema on Twitter / X
This is why we need open benchmarks for local AI.Otherwise it turns into tribalism and name calling.We will be publishing the largest database of open benchmarks for local AI, tested on 1,000+ real hardware setups. Every device, every interconnect, different… https://t.co/ZsU3PCdSsZ— Alex Cheema (@alexocheema) March 9, 2026
Mirai Labs: Frontier On-Device AI Lab
Models, runtime & infrastructure to make on-device AI interactive, ambient & continuous.

Optimizing On-Device Inference for Apple Silicon
A custom local engine that improves prefill and decode throughput

ZETIC | On-Device AI for Everything - for any model, on any device, in any framework
Built by ex-Qualcomm AI Engineer. Automate on-device AI deployment with full NPU optimization. Benchmark on 100+ physical devices and ship in hours with just 3 lines of code.

Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

Open Source AI Inference Benchmark | InferenceX
Compare AI inference performance across GPUs and frameworks. Real benchmarks on NVIDIA GB200, B200, AMD MI355X, and more. Free, open-source, continuously updated.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Introducing On device AI capabilities inside the Craft Assistant
A first look at the future of on-device AI models integrated into productivity tools

Google for Developers Blog - News about Web, Mobile, AI and Cloud
LiteRT is the universal framework for on-device AI. The production stack delivers 1.4x faster cross-platform GPU performance, streamlined NPU acceleration, and superior GenAI support for open models like Gemma.

Wafer - Ship the fastest inference in the world
Autonomous AI agents that profile, diagnose, and optimize GPU inference across your entire stack — from kernels to models to production pipelines.

Open-Source Agentic Inference Benchmark | InferenceX
Compare AgentX, InferenceX's long-context, multi-turn coding scenario, with fixed-sequence AI inference across chips and frameworks. Public NVIDIA and AMD runs update when configurations change.