







Low-latency AI engine for mobile devices & wearables
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Melange | On-device AI for Mobile
Select. Benchmark. Deploy | End-to-end on-device AI deployment tool for mobile devs.
Introducing Pipette: A benchmarking suite for on-device intelligence — Blog
Meet Pipette, an open-source platform for reproducible on-device AI benchmarks across models, quantization, runtimes and hardware.
Scott Jenson – Exploring the world beyond mobile
Fast AI requires slow thinking The current Silicon Valley flex is trading notes on your favorite new AI tools over lunch. Each week brings another one to explore. Some of these tools are very impressive; I’ve been able to reply to someone with an alternative UI design in less than a minute, and they were […]
AI-Native Cloud | DigitalOcean
Run AI products in production with a unified stack for agents, inference, and cloud—built for control, performance, and economics at scale.
Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

Gabber - Build Realtime AI Apps that can see, hear, and speak
Low-latency inference for VLM, TTS, and STT with orchestration for making realtime apps.

ZETIC | On-Device AI for Everything - for any model, on any device, in any framework
Built by ex-Qualcomm AI Engineer. Automate on-device AI deployment with full NPU optimization. Benchmark on 100+ physical devices and ship in hours with just 3 lines of code.

Locally AI - Run AI models locally on your iPhone, iPad, and Mac.
Run Llama, Gemma, Qwen, DeepSeek, and more on your iPhone, iPad, and Mac. Optimized for Apple Silicon. Offline. Private.

Alex Cheema on Twitter / X
This is why we need open benchmarks for local AI.Otherwise it turns into tribalism and name calling.We will be publishing the largest database of open benchmarks for local AI, tested on 1,000+ real hardware setups. Every device, every interconnect, different… https://t.co/ZsU3PCdSsZ— Alex Cheema (@alexocheema) March 9, 2026
TheStage AI – Faster, Cheaper AI Inference
Accelerate models on NVIDIA & edge. Full guides for setup, optimization & deploy. ANNA, QLIP, Elastic Models, CLI & API. Built for AI teams & devs.

Part 3: iPhone Hardware and How It Powers On-Device AI
By Mirai Labs, frontier on-device AI lab. Building the models, inference runtime, and quantization stack from the device constraint up.

Google for Developers Blog - News about Web, Mobile, AI and Cloud
LiteRT is the universal framework for on-device AI. The production stack delivers 1.4x faster cross-platform GPU performance, streamlined NPU acceleration, and superior GenAI support for open models like Gemma.

google-ai-edge/LiteRT
LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
Latency optimization | OpenAI API
Improve latency across a wide variety of LLM-related use cases.
