







👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025
BerriAI/litellm
Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]
Osaurus — Own Your AI on Apple Silicon
Own your AI: local-first agents with memory, tools, and identity on Apple Silicon. Offline, open source, and API-compatible with OpenAI, Anthropic, and Ollama.

Osaurus — Own Your AI on Apple Silicon
Own your AI: local-first agents with memory, tools, and identity on Apple Silicon. Offline, open source, and API-compatible with OpenAI, Anthropic, and Ollama.

OpenAI Developers on Twitter / X
Today we’re announcing Open Responses: an open-source spec for building multi-provider, interoperable LLM interfaces built on top of the original OpenAI Responses API.✅ Multi-provider by default✅ Useful for real-world workflows✅ Extensible without fragmentationBuild… pic.twitter.com/SJiBFx1BOF— OpenAI Developers (@OpenAIDevs) January 15, 2026
Ollama is now powered by MLX on Apple Silicon in preview· Ollama Blog
Today, we're previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple's machine learning framework.

raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Latency optimization | OpenAI API
Improve latency across a wide variety of LLM-related use cases.

Issue · BerriAI/litellm
Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthr...
OpenAI on Twitter / X
Our open models are here.Both of them.https://t.co/9tFxefOXcg— OpenAI (@OpenAI) August 5, 2025
OpenAI gpt-oss· Ollama Blog
Ollama partners with OpenAI to bring gpt-oss to Ollama and its community.

Build Bigger With Small Ai: Running Small Models Locally
Stream29/ProxyAsLocalModel
Proxy remote LLM API as Ollama and LM Studio, for using them in JetBrains AI Assistant
Geek Lite on Twitter / X
在 Apple Silicon Mac 上本地运行 LLM 推理服务,提供比 Ollama 和 llama.cpp 更快的 OpenAI 兼容 API,同时原生支持工具调用和提示缓存。https://t.co/IXet9GW7x0Rapid-MLX 用 Apple 自家的 MLX 框架做推理,搭了个 FastAPI 服务跑 OpenAI 兼容 API。在 Apple Silicon 上比 Ollama 快 2-4 倍,靠… pic.twitter.com/d9QUbAWJSZ— Geek Lite (@QingQ77) May 3, 2026

open-slopware
Free/Open Source Software choosing to use and/or support LLM usage/AI, as well as alternatives and tips to requesting better policies or forking.