







Democratizing AI compute through heterogeneous device orchestration
Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

Introducing Pipette: A benchmarking suite for on-device intelligence — Blog
Meet Pipette, an open-source platform for reproducible on-device AI benchmarks across models, quantization, runtimes and hardware.
Flyte | One Platform for Your AI Orchestration Needs
Dynamic, resilient AI orchestration. 80M+ downloads.

google-ai-edge/LiteRT
LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
SambaNova | The Fastest AI Inference Platform
Discover SambaNova - the complete AI platform delivering the fastest AI inference, fine-tuning, and scalable solutions for agentic AI easily integrated into existing data center infrastructures.

Lilypad-Tech/lilypad
Run AI workloads easily in a decentralized GPU network. https://www.youtube.com/watch?v=yQnB2Yxia4Y
TurboQuant: Redefining AI efficiency with extreme compression
Amir Zandieh, Research Scientist, and Vahab Mirrokni, VP and Google Fellow, Google Research

Melange | On-device AI for Mobile
Select. Benchmark. Deploy | End-to-end on-device AI deployment tool for mobile devs.
The Universal Execution Layer for AI
Optimize any AI model on any engine, across all hardware. Dria’s topology-aware compiler and peer-to-peer runtime merge CPUs, GPUs, NPUs & chiplets into one fabric—maximising utilisation, cutting inference cost and ending vendor lock-in.

ZETIC | On-Device AI for Everything - for any model, on any device, in any framework
Built by ex-Qualcomm AI Engineer. Automate on-device AI deployment with full NPU optimization. Benchmark on 100+ physical devices and ship in hours with just 3 lines of code.

On-device intelligence for every product
We're building a frontier lab for on-device AI. Small, specialized models for audio, vision, and text, faster than the cloud and free.

Teaching AI to Optimize AI Models for Edge Deployment
How our agent, Möbius automated a Core ML port in ~12 h (vs. 2 weeks), hit 0.99998 parity, and made it 3.5× faster, while staying on the CPU.

Alex Cheema on Twitter / X
This is why we need open benchmarks for local AI.Otherwise it turns into tribalism and name calling.We will be publishing the largest database of open benchmarks for local AI, tested on 1,000+ real hardware setups. Every device, every interconnect, different… https://t.co/ZsU3PCdSsZ— Alex Cheema (@alexocheema) March 9, 2026
AI-Native Cloud | DigitalOcean
Run AI products in production with a unified stack for agents, inference, and cloud—built for control, performance, and economics at scale.
Wafer - Ship the fastest inference in the world
Autonomous AI agents that profile, diagnose, and optimize GPU inference across your entire stack — from kernels to models to production pipelines.
