







An interactive demo for developers to try the latest text-to-speech model in the OpenAI API
Realtime and audio | OpenAI API
Learn which realtime and audio guide to use for each speech application.

How OpenAI delivers low-latency voice AI at scale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.

DeepL AI Platform: Translation, Voice & API
Explore our AI suite and get more done: Translate speech, text, and media, or integrate the DeepL API.
Building Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAI
Simon Willison on Twitter / X
I think it's non-obvious to many people that the OpenAI voice mode runs on a much older, much weaker model - it feels like the AI that you can talk to should be the smartest AI but it really isn't https://t.co/bZ0Qqx9Sa9— Simon Willison (@simonw) April 10, 2026
Introducing Whisper – an open source voice note taking app! Record voice notes and transcribe them into lists, blogs, & more with AI. 100% free & open source. https://t.co/UZWGkUDJ6d
Introducing Whisper – an open source voice note taking app!Record voice notes and transcribe them into lists, blogs, & more with AI.100% free & open source. pic.twitter.com/UZWGkUDJ6d— Hassan (@nutlope) July 22, 2025
OpenAI Harmony Response Format
The gpt-oss models were trained on the harmony response format for defining conversation structures, generating reasoning output and structu

\robotoslablightdots.tts Technical Report
Text-to-speech (TTS) systems have largely solved intelligibility on standard read-speech benchmarks. What users expect from a modern system is broader: expressive and controllable output, real-time synthesis, and coverage of neutral reading, emotional dialogue, paralinguistic events, singing, and general audio. Current systems pursue this goal along three roughly distinct technical routes, and each route has its own unresolved problem.
GPT-Live System Card - OpenAI Deployment Safety Hub
GPT-Live-1 and GPT-Live-1 mini are a new generation of voice models designed to make conversations with AI feel more natural and intelligent.

Function calling | OpenAI API
Learn how function calling enables large language models to connect to external data and systems.

OpenAI's WebRTC Problem - Media over QUIC
Media over QUIC: There are ways to do voice AI without being traumatized by WebRTC.


Introducing GPT-Live
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.

Luozhu on Twitter / X
I taught a speech model to understand context in conversation. This is what happenedIt adjusts voice and tone to express urgency, comfort, understanding from the dialogue. Just like a real human being520M model. Runs locally on consumer devicesHow this is achieved 🧵 pic.twitter.com/oMgSh2TKbO— Luozhu (@LuozhuZhang) February 27, 2026
Voice AI & Voice Agents | An Illustrated Primer
A comprehensive guide to voice AI in 2026
