







Avanika Narayan, Dan Biderman, and Sabri Eyuboglu from Christopher Ré's Stanford Hazy Research lab, along with Avner May, Scott Linderman, James Zou, have developed a way to shift a substantial portion of LLM workloads to consumer devices by having small on-device models (such as Llama 3.2 with Ollama) collaborate with larger models in the cloud (such as GPT-4o).
Extensible Software in the age of LLMs | Jeremy Morrell
Solid core + capability-based sandboxes + LLMs = Users with superpowers

Secure Minions: private collaboration between Ollama and frontier models· Ollama Blog
Secure Minions is a secure protocol built by Stanford's Hazy Research lab to allow encrypted local-remote communication.

Build Bigger With Small Ai: Running Small Models Locally
Alex Cheema on Twitter / X
.@karpathy shouted out my work on @exolabs at @ycombinator AI SUS!“we use LLMs similarly to mainframes in the ‘70s - compute is timeshared by having a slice in the batch dimension. models will compress over time, and with this we’ll be able to run more on-device” pic.twitter.com/UKSrquQGAL— Alex Cheema (@alexocheema) June 18, 2025

Featherless - Serverless LLM Hosting
Freedom to reliably deploy any open model effortlessly.

Home - NobodyWho
NobodyWho is an inference engine that lets you run LLMs locally on any device
LM Link: Access models on your powerful devices you own, as if they were local
Tailscale and LM Studio partner to provide encrypted access to remote LLMs on hardware you own.

The past, present, and future of local-first - Martin Kleppmann (Local-First Conf)
The past, present, and future of local-first - Martin Kleppmann (Local-First Conf)
Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

Mesh-LLM/mesh-llm
Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.
mzau/broke-cluster
A Poor Man's Apple Silicon LLM Cluster — tuned for MLX, scalable without shame.
Benchmarks | EXO
Transparent benchmarks for LLMs tested on real hardware. Coming soon.
qualcomm/nexa-sdk
Run frontier LLMs and VLMs with day-0 model support across GPU, NPU, and CPU, with comprehensive runtime coverage for PC (Python/C++), mobile (Android & iOS), and Linux/IoT (Arm64 & x86 Docker). Supporting OpenAI GPT-OSS, IBM Granite-4, Qwen-3-VL, Gemma-3n, Ministral-3, and more.