







Hundreds of models & providers. One command to find what runs on your hardware.
clem 🤗 on Twitter / X
We just released an hf CLI extension to detect the best model/quant for a user's hardware and then spins up a local coding agent. Time to go local/private/free/fast for your agents thanks to open-source! pic.twitter.com/LcVJzGCqWx— clem 🤗 (@ClementDelangue) March 17, 2026

apple-silicon-llm-bench/results/complete_results.html at main · AlexHiesch/apple-silicon-llm-bench
Systematic LLM inference benchmark for Apple Silicon: 8 backends, 7 models, 791 measurements - AlexHiesch/apple-silicon-llm-bench
Benchmarks | EXO
Transparent benchmarks for LLMs tested on real hardware. Coming soon.
qualcomm/nexa-sdk
Run frontier LLMs and VLMs with day-0 model support across GPU, NPU, and CPU, with comprehensive runtime coverage for PC (Python/C++), mobile (Android & iOS), and Linux/IoT (Arm64 & x86 Docker). Supporting OpenAI GPT-OSS, IBM Granite-4, Qwen-3-VL, Gemma-3n, Ministral-3, and more.
Running local models on an M4 with 24GB memory | jola.dev
Experiments with getting usable outputs out of local models on a standard Macbook

How to check the detail specifications of my ASUS product? | Official Support | ASUS Canada
For specific operations, you can refer to the explanation in the [Product Footprint]section here.
OpenRouter
The unified interface for LLMs. Find the best models & prices for your prompts
On-Device LLM Throughput Calculator - a Hugging Face Space by FL33TW00D-HF
This tool estimates and visualizes the throughput of Large Language Models on devices with memory bandwidth constraints. Users input device and model configurations, and the tool generates a plot s...
Can You Run This LLM? VRAM Calculator (Nvidia GPU and Apple Silicon)
Calculate the VRAM required to run any large language model.

Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

Home - NobodyWho
NobodyWho is an inference engine that lets you run LLMs locally on any device
open-slopware
Free/Open Source Software choosing to use and/or support LLM usage/AI, as well as alternatives and tips to requesting better policies or forking.
InferenceMAX™: Open Source Inference Benchmarking
NVIDIA GB200 NVL72, AMD MI355X, Throughput Token per GPU, Latency Tok/s/user, Perf per Dollar, Tokens per Provisioned Megawatt, DeepSeek R1 670B, GPTOSS 120B, Llama3 70B

Compiling Models to Megakernels
Fine-grained synchronization, deep pipelines, and zero kernel launch overheads, automatically.
