







OpenAI's OSS model possible breakdown:1. 120B MoE 5B active + 20B text only2. Trained with Float4 maybe Blackwell chips3. SwiGLU clip (-7,7) like ReLU64. 128K context via YaRN from 4K5. Sliding window 128 + attention sinks6. Llama/Mixtral arch + biasesDetails:1. 120B MoE… https://t.co/bMFp3Z6Gs5 pic.twitter.com/1NFO4utPqr— Daniel Han (@danielhanchen) August 1, 2025
Release v0.11.0 · ollama/ollama
Welcome OpenAI's gpt-oss models Ollama partners with OpenAI to bring its latest state-of-the-art open weight models to Ollama. The two models, 20B and 120B, bring a whole new local chat experie...
The OpenAI Open weight model might be 120B
726 votes, 163 comments. The person who "leaked" this model is from the openai (HF) organization So as expected, it's not gonna be something you can…
gpt-oss:120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

Georgi Gerganov on Twitter / X
gpt-oss is a great modelIMO OpenAI showed us the blueprint for winning local AI:- Interleaved SWA- Small head sizes in the attention- Attention sinks- Mixture of Experts FFN- 4-bit trainingAll of these parts combined together result in the best architecture suitable for…— Georgi Gerganov (@ggerganov) August 28, 2025
OpenAI Developers on Twitter / X
Today we’re announcing Open Responses: an open-source spec for building multi-provider, interoperable LLM interfaces built on top of the original OpenAI Responses API.✅ Multi-provider by default✅ Useful for real-world workflows✅ Extensible without fragmentationBuild… pic.twitter.com/SJiBFx1BOF— OpenAI Developers (@OpenAIDevs) January 15, 2026
OpenAI on Twitter / X
Our open models are here.Both of them.https://t.co/9tFxefOXcg— OpenAI (@OpenAI) August 5, 2025

gpt-oss:20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

OpenAI gpt-oss· Ollama Blog
Ollama partners with OpenAI to bring gpt-oss to Ollama and its community.

Rohan Paul on Twitter / X
👨🔧 Github: Native, Apple Silicon–only local LLM server. Similar to Ollama, but built on Apple's MLX- OpenAI API compatible, - Ollama‑compatible- OpenAI‑style tools + tool_choice, with tool_calls parsing and streaming deltasgithub. com/dinoki-ai/osaurus pic.twitter.com/G0NbWnEJcQ— Rohan Paul (@rohanpaul_ai) September 3, 2025

Hatice Ozen on Twitter / X
PSA: @OpenAI is putting the Open back in OpenAI and @GroqInc has Day 0 support. 🤗GPT-OSS 20B and 120B, hybrid-reasoning models with built-in browser search and code execution are now live for instant inference.P.S. We've also launched OpenAI Responses API compatibility. pic.twitter.com/CK7StvMSpr— Hatice Ozen (@ozenhati) August 5, 2025
Ivan Fioravanti ᯅ on Twitter / X
"We are releasing Open Source implementations for CoreAILanguageModel and MLXLanguageModel for running a myriad of local models on the Apple Neural Engine or your Mac's GPU" 👀 From #WWDC26: What’s new in the Foundation Models framework video: https://t.co/1NtKWYhNRs pic.twitter.com/HVtBsr3tjL— Ivan Fioravanti ᯅ (@ivanfioravanti) June 9, 2026
OpenAI partners with Cerebras
OpenAI partners with Cerebras to add 750MW of high-speed AI compute, reducing inference latency and making ChatGPT faster for real-time AI workloads.

How OpenAI delivers low-latency voice AI at scale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.

OpenAI Codex with Ollama· Ollama Blog
Open models can be used with OpenAI's Codex CLI through Ollama. Codex can read, modify, and execute code in your working directory using models such as gpt-oss:20b, gpt-oss:120b, or other open-weight alternatives.
