







MiniCPM5: SOTA on-device LLMs, small yet powerful.
Minions: where local and cloud LLMs meet· Ollama Blog
Avanika Narayan, Dan Biderman, and Sabri Eyuboglu from Christopher Ré's Stanford Hazy Research lab, along with Avner May, Scott Linderman, James Zou, have developed a way to shift a substantial portion of LLM workloads to consumer devices by having small on-device models (such as Llama 3.2 with Ollama) collaborate with larger models in the cloud (such as GPT-4o).

Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

Needle 2 - The 14 MB Agentic LLM for Tiny Devices | Cactus
An open 45M-parameter model for tool calling, device use, and structured extraction. Needle 2 runs as a 14 MB binary in 28 MB of session RAM.
Benchmarks | EXO
Transparent benchmarks for LLMs tested on real hardware. Coming soon.
microsandbox - Every agent deserves its own machine
Run lightweight microVMs locally. Programmable networking, custom filesystems, secrets that never leak.
mzau/broke-cluster
A Poor Man's Apple Silicon LLM Cluster — tuned for MLX, scalable without shame.
ARahim3/mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
ARahim3/mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
PL Series|Mini PCs|ASUS Global
Ultraslim and reliable Mini PCs with rich I/O connectivity, designed for vertical markets

Home - NobodyWho
NobodyWho is an inference engine that lets you run LLMs locally on any device
Build Bigger With Small Ai: Running Small Models Locally
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
Added a new device to my @tiles.run cluster. Welcome to the lineup, M5 Pro with the 32 GB/1 TB spec.