







Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B c...
noonghunna/club-3090
Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
noname on Twitter / X
Upto 1100 tps on RTX 3090x2 for Diffusion Gemma 4 26B.Unleash this mini monster on your gpus now!If you are running nvidia gpus locally, come grab the recipe at club-3090. https://t.co/qKuFcgu1llP.S. a ⭐️ on Github is much appreciated.@googlegemma @vllm_project— noname (@malikwas1f) June 11, 2026
crawshaw - 2026-02-08
I wrote up my experiences programming with LLMs a bit over a year ago, and updated it for the world of agents eight months ago. A lot has changed since then, so here is an update.
Joel - coffee/acc on Twitter / X
OKAY - it seemed like DFlash would be the clear winner.But it appears there have been some improvements with MTP.With MTP + split-mode = tensor, Qwen3.6-27B gets over 120 tokens/second on dual RTX 3090s (note I am running PCIE x16 on both, I don't have an NVLink bridge).… https://t.co/AQmqTJrof7 pic.twitter.com/RQSPVsNwMO— Joel - coffee/acc (@JoelDeTeves) June 29, 2026

at://userandagents.org/network.cosmik.card/Kffy9lJlvKgl-KCjGSTHymU4
OpenRouter
The unified interface for LLMs. Find the best models & prices for your prompts
Release v0.29.0 · ml-explore/mlx
Highlights Support for mxfp4 quantization (Metal, CPU) More performance improvements, bug fixes, features in CUDA backend mx.distributed supports NCCL back-end for CUDA What's Changed [CUDA]...
Qwen3.6 27B — Qwen/Qwen3.6-27B rental
Free · 32 concurrent · hosted by selimaktas on LocalMaxxing.
Xuan-Son Nguyen on Twitter / X
Firefox is open-source on Github, and they experimented with @ggml_org llama.cpp in WASM 👀Wondering what they are cooking 🧑🍳 pic.twitter.com/mMtVZo0gc2— Xuan-Son Nguyen (@ngxson) May 13, 2025

Victor M on Twitter / X
llama.cpp UI now has MCP support 🔥Working super well, make sure you are up to date:```brew install llama.cpp```then```llama-server --webui-mcp-proxy``` pic.twitter.com/r8JXwpbfHq— Victor M (@victormustar) March 9, 2026

Just pushed a bunch of cards to my @semble.so collection via the API from my bookmarks! thank you opus and thank you @cosmik.network team for the API! tab sanity 🤝 contributing to the knowledge commons 🎉