







Local LLMs, Self-Hosting, and Hardware
Home - NobodyWho
NobodyWho is an inference engine that lets you run LLMs locally on any device
Featherless - Serverless LLM Hosting
Freedom to reliably deploy any open model effortlessly.

How to Serve Local LLMs Anywhere: Secure Remote Access with Cloudflare and Unsloth | Unsloth Documentation
Unsloth is an open-source project that allows you to train and run LLMs locally and with Cloudflare tunnel, you can access Unsloth from your mobile device, share access to a friend or coworker, host Unsloth on a server such as Google Colab, AWS, or even a personal server.

LM Link: Access models on your powerful devices you own, as if they were local
Tailscale and LM Studio partner to provide encrypted access to remote LLMs on hardware you own.

sauravpanda/BrowserAI
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser
Minions: where local and cloud LLMs meet· Ollama Blog
Avanika Narayan, Dan Biderman, and Sabri Eyuboglu from Christopher Ré's Stanford Hazy Research lab, along with Avner May, Scott Linderman, James Zou, have developed a way to shift a substantial portion of LLM workloads to consumer devices by having small on-device models (such as Llama 3.2 with Ollama) collaborate with larger models in the cloud (such as GPT-4o).

Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

Reuse your existing hardware to run LLMs privately and securely.
Trellis lets you run large language models on your organization's compute. Scale and data privacy, choose both.

Open Responses
This is the standardization effort I've most wanted in the world of LLMs: a vendor-neutral specification for the JSON API that clients can use to talk to hosted LLMs. Open …
Run LLMs locally on your Mac · mlx-optiq
Quantize, fine-tune and serve LLMs locally on Apple Silicon. MLX-native, no PyTorch, no cloud. On PyPI.

Hoping @letta.com is still on everyones Radar. They somewhat recently added & then massively expanded support for a local mode. w/ app-server mode it seems like virtually all capabilities they advertise can now be self hosted on your own LLM. The product is insanely good! So happy they did this.