







Train the smallest LM you can that fits in 16MB. Best model wins!
Overview - GroqDocs
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.

OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.

Latency optimization | OpenAI API
Improve latency across a wide variety of LLM-related use cases.

Build Bigger With Small Ai: Running Small Models Locally
Needle 2 - The 14 MB Agentic LLM for Tiny Devices | Cactus
An open 45M-parameter model for tool calling, device use, and structured extraction. Needle 2 runs as a 14 MB binary in 28 MB of session RAM.
Liquid AI on Twitter / X
Today, we release LFM2.5-350M. Agentic loops at 350M parameters.A 350M model trained for reliable data extraction and tool use, where models at this scale typically struggle.<500MB when quantized, built for environments where compute, memory, and latency are constrained.🧵 pic.twitter.com/zZPKzcCwH9— Liquid AI (@liquidai) March 31, 2026

Daniel Han on Twitter / X
OpenAI's OSS model possible breakdown:1. 120B MoE 5B active + 20B text only2. Trained with Float4 maybe Blackwell chips3. SwiGLU clip (-7,7) like ReLU64. 128K context via YaRN from 4K5. Sliding window 128 + attention sinks6. Llama/Mixtral arch + biasesDetails:1. 120B MoE… https://t.co/bMFp3Z6Gs5 pic.twitter.com/1NFO4utPqr— Daniel Han (@danielhanchen) August 1, 2025

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show | TechCrunch
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

On-Device LLM Throughput Calculator - a Hugging Face Space by FL33TW00D-HF
This tool estimates and visualizes the throughput of Large Language Models on devices with memory bandwidth constraints. Users input device and model configurations, and the tool generates a plot s...
OpenAI Codex with Ollama· Ollama Blog
Open models can be used with OpenAI's Codex CLI through Ollama. Codex can read, modify, and execute code in your working directory using models such as gpt-oss:20b, gpt-oss:120b, or other open-weight alternatives.

OpenMemory - AI Memory MCP Server for Coding Agents | Mem0
With OpenMemory, add persistent, project-aware memory to Cursor, Windsurf, and VS Code agents. Store preferences, patterns, and context that get retrieved automatically.

Run DeepSeek-R1 Dynamic 1.58-bit
DeepSeek R-1 is the most powerful open-source reasoning model that performs on par with OpenAI's o1 model. Run the 1.58-bit Dynamic GGUF version by Unsloth.

OpenCode Go | Low cost coding models for everyone
Go starts at $5 for your first month, then $10/month, with generous 5-hour request limits for GLM-5.1, GLM-5, Kimi K2.5, Kimi K2.6, MiMo-V2.5-Pro, MiMo-V2.5, Qwen3.7 Max, Qwen3.6 Plus, MiniMax M2.5, MiniMax M2.7, MiniMax M3, DeepSeek V4 Pro, and DeepSeek V4 Flash.

How I fit a 28.9M LLM on an ESP32-S3 (~9 tok/s, fully on-chip)
1.2K votes, 69 comments. I wanted to see how big a language model I could actually run on an ESP32. Not the 260K-param TinyStories model that's been…
If you ever wanted to know how big the Bluesky/atproto network is (records only, no blobs/videos/etc), there is a new, super useful tool in town: -> jetstream.us-east.bsky.network/status?tab=segments ~ 1.7 TB compressed / 8 TB uncompressed