







Simple model memory requirements calculator for GGUF
On-Device LLM Throughput Calculator - a Hugging Face Space by FL33TW00D-HF
This tool estimates and visualizes the throughput of Large Language Models on devices with memory bandwidth constraints. Users input device and model configurations, and the tool generates a plot s...
Thireus/GGUF-Tool-Suite
Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain will create a GGUF recipe tuned to your hardware within seconds — flexible model sizing and lowest achievable perplexity/kld for GGUF enthusiasts seeking precise and automated dynamic quant production.
Reverse-engineering GGUF | Post-Training Quantization
Magnitude
The best way to code with open models. Magnitude has 9 repositories available. Follow their code on GitHub.
Can You Run This LLM? VRAM Calculator (Nvidia GPU and Apple Silicon)
Calculate the VRAM required to run any large language model.


Model Release Notes | OpenAI Help Center
We’re beginning the rollout of GPT-5.6 Sol in ChatGPT, our flagship reasoning model for complex work across coding, research, science, cybersecurity, computer use, and design.GPT-5.6 Sol is rolling out to eligible paid ChatGPT plans. Free, Go, and logged-out users are not included. Availability may vary during rollout, and managed-workspace access can depend on administrator settings. Availability for other GPT-5.6 family models varies by product and plan; check the model picker or current rate card in the product you use.

OpenAI Codex with Ollama· Ollama Blog
Open models can be used with OpenAI's Codex CLI through Ollama. Codex can read, modify, and execute code in your working directory using models such as gpt-oss:20b, gpt-oss:120b, or other open-weight alternatives.

Release v0.11.0 · ollama/ollama
Welcome OpenAI's gpt-oss models Ollama partners with OpenAI to bring its latest state-of-the-art open weight models to Ollama. The two models, 20B and 120B, bring a whole new local chat experie...
Awni Hannun on Twitter / X
It's very cool that Apple shipped a 20B parameter on-device. You can't put 20B parameters in RAM at any reasonable precision. To make it work they are using pretty exotic architecture by today's standards.A small model predicts from the query (or prompt) which experts to load… pic.twitter.com/Zhe5HcbGuL— Awni Hannun (@awnihannun) June 9, 2026

Overview - GroqDocs
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.

Qalculate! - the ultimate desktop calculator
Modelfile Reference - Ollama English Documentation
ollama 的中英文文档,中文文档由 llamafactory.cn 翻译