







Media over QUIC: There are ways to do voice AI without being traumatized by WebRTC.
How OpenAI delivers low-latency voice AI at scale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.

OpenAI.fm
An interactive demo for developers to try the latest text-to-speech model in the OpenAI API

Realtime and audio | OpenAI API
Learn which realtime and audio guide to use for each speech application.

Media over QUIC
Media over QUIC is a new live media protocol designed for simplicity and scale. It uses new browser technologies like WebTransport and WebCodecs to deliver media with latency that rivals WebRTC.
Simon Willison on Twitter / X
I think it's non-obvious to many people that the OpenAI voice mode runs on a much older, much weaker model - it feels like the AI that you can talk to should be the smartest AI but it really isn't https://t.co/bZ0Qqx9Sa9— Simon Willison (@simonw) April 10, 2026
Voice AI & Voice Agents | An Illustrated Primer
A comprehensive guide to voice AI in 2026

DeepL AI Platform: Translation, Voice & API
Explore our AI suite and get more done: Translate speech, text, and media, or integrate the DeepL API.
Building Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAI
Pipecat Cloud: Enterprise Voice Agents Built On Open Source - Kwindla Hultman Kramer, Daily
GPT-Live System Card - OpenAI Deployment Safety Hub
GPT-Live-1 and GPT-Live-1 mini are a new generation of voice models designed to make conversations with AI feel more natural and intelligent.

Introducing GPT-Live
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.


Why ChatGPT Keeps Interrupting You — Dr. Tom Shapland, LiveKit
Introducing Whisper – an open source voice note taking app! Record voice notes and transcribe them into lists, blogs, & more with AI. 100% free & open source. https://t.co/UZWGkUDJ6d
Introducing Whisper – an open source voice note taking app!Record voice notes and transcribe them into lists, blogs, & more with AI.100% free & open source. pic.twitter.com/UZWGkUDJ6d— Hassan (@nutlope) July 22, 2025
Coqui TTS & XTTS V2: AI Text to Speech in 8 Languages
Experience natural speech synthesis with Coqui TTS and XTTS V2 technology. Features voice cloning and support for 8 languages.
I think of all of the AI / ML / CS tech out there, speech generation freaks me out the most.
Opensourcing TADA: Fast, Reliable Speech Generation Through Text-Acoustic Synchronization
www.hume.ai