







OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets
OpenAI says its Jalapeño chip can power faster AI responses than the competition
OpenAI still isn’t giving up Nvidia chips, though.

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show | TechCrunch
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

Jalapeño’s first results show industry-leading speed and efficiency in AI inference
Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.

How will OpenAI compete? — Benedict Evans
OpenAI has some big questions. It doesn’t have unique tech. It has a big user base, but with limited engagement and stickiness and no network effect. The incumbents have matched the tech and are leveraging their product and distribution. And a lot of the value and leverage will come from new experie

OpenAI partners with Cerebras
OpenAI partners with Cerebras to add 750MW of high-speed AI compute, reducing inference latency and making ChatGPT faster for real-time AI workloads.

OpenAI gpt-oss· Ollama Blog
Ollama partners with OpenAI to bring gpt-oss to Ollama and its community.

Teaching Everyone to Fish for Tokens
Nvidia wants you building your own model, not buying from Anthropic/OpenAI.


OpenAI buys local-first, plus Claude and MedGemma health updates
trust is hard to win, easy to lose, and really expensive to acquire.

nvidia/nemotron-3.5-asr-streaming-0.6b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
InferenceMAX™: Open Source Inference Benchmarking
NVIDIA GB200 NVL72, AMD MI355X, Throughput Token per GPU, Latency Tok/s/user, Perf per Dollar, Tokens per Provisioned Megawatt, DeepSeek R1 670B, GPTOSS 120B, Llama3 70B

NVIDIA Shatters MoE AI Performance Records With a Massive 10x Leap on GB200 'Blackwell' NVL72 Servers, Fueled by Co-Design Breakthroughs
Scaling performance on MoE AI models is one of the industry constraints, but it appears that NVIDIA has managed to make a breakthrough.
