







Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

BerriAI/litellm
Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]
Using Ollama with Kilo Code | Run Local Models
Run local AI models with Ollama in Kilo Code for offline, private coding. Setup guide for VS Code and the CLI.
Issue · BerriAI/litellm
Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthr...
freestylefly/wesight
Open-source desktop AI agent workspace with one-click Claude Code, Codex, OpenClaw, Hermes Agent setup and custom LLM model routing.
Build Bigger With Small Ai: Running Small Models Locally
Streaming responses with tool calling· Ollama Blog
Ollama now supports streaming responses with tool calling. This enables all chat applications to stream content and also call tools in real time.

Benjamin Marie (@bnjmnmarie)
A new alternative to Ollama: You can now run models directly through Unsloth (with Docker): https://docs.unsloth.ai/models/how-to-run-llms-with-docker It supports the same models as llama.cpp, which I guess means it runs on llama.cpp… But this way you don’t need to set up anything, if you already have Docker installed.

OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.

Open Responses with local models via LM Studio
Update to LM Studio 0.3.39 for Open Responses support

Introducing any-llm: A unified API to access any LLM provider
When it comes to using LLMs, it’s not always a question of which model to use: it’s also a matter of choosing who provides the LLM and where it is deployed. Today, we announce the release of any-llm, a Python library that provides a simple unified interface to access the most popular providers.

Mesh LLM: distributed AI computing on iroh
How Mesh LLM pools existing GPU resources across machines into a single OpenAI-compatible API, built on iroh.
Minions: where local and cloud LLMs meet· Ollama Blog
Avanika Narayan, Dan Biderman, and Sabri Eyuboglu from Christopher Ré's Stanford Hazy Research lab, along with Avner May, Scott Linderman, James Zou, have developed a way to shift a substantial portion of LLM workloads to consumer devices by having small on-device models (such as Llama 3.2 with Ollama) collaborate with larger models in the cloud (such as GPT-4o).

jjang-ai/mlxstudio
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
LM Studio on Twitter / X
Introducing LM Link ✨ Connect to remote instances of LM Studio, securely.🔐 End-to-end encrypted📡 Load models locally, use them on the go🖥️ Use local devices, LLM rigs, or cloud VMsLaunching in partnership with @TailscaleTry it now: https://t.co/2iB4mwjfxy— LM Studio (@lmstudio) February 25, 2026
distil labs — Replace LLMs with Custom Small Language Models
Train and deploy custom small language models that are faster, cheaper, and just as accurate as LLMs.
Proxy remote LLM API as Ollama and LM Studio, for using them in JetBrains AI Assistant