7 ms·
How OpenAI delivers low-latency voice AI at scale
- AIorNot 5mo agoso is the answer WebRTC + Kubernetes
- anzerarkin 5mo agoI hate the voice ai though, it's so much dumber
- NikolaNovak 5mo agoFwiw - I found the advanced AI voice feature to be actually detrimental. It's good if you just want a single sentence answer. I've turned it off though when I want a more detailed, structured, considered answer.
- drusepth 5mo agoInterestingly, that kind of parallels the real world too: if you want a quick and high level answer, talk to someone in person; if you want something detailed and info-dense, get them to write it down.
- NikolaNovak 5mo agoTurning advanced voice still leaves "regular" voice interaction which are actually (for me:) much much better - it's just the regular response, verbalized :). Voice quality isn't worse, it just doesn't try to summarize in one casual sentence. (I still hate that the voice is getting more and more "natural" - the umms and ahhs and weird pauses)
- hadlock 5mo agoThe AI attached to their voice chat is running a completely different model. Ask any question and you quickly realize it is completely, unapologetically lobotomized. If you want to talk to it about how you feel after your gf/bf broke up with you, it is fine. If you want to ask it something about tunneling machines and how tunneling through different types of rock impact engineering decisions, it is going to skim the first four sentences of some blog article and then defend whatever hill it has chosen until it dies, regardless of what the larger body of work on the topic says. OpenAI's voice chat being so bad and being totally divorced from their SOTA models is largely why I cancelled my subscription. I am tempted to wire up piper/whisper and the OAI api to get back what I actually want/need. But today you cannot have a conversation about engineering questions and get anything close to factually reliable answers out of it.
- brett-jackson 5mo agoI used to use it all the time until about a year ago or so. Its responses are full of filler and the safeguards are really overbearing. It often will just give wrong answers in a way that GPT-5.x does not. I once asked it why a particular celebrity was canceled and it refused to tell me because it may harm me to know what they said!
- cdrnsf 5mo agoIt's missing the part where they explain how they obtained the training data for their voice AI.
- thimabi 5mo ago> Voice AI only feels natural if conversation moves at the speed of speech […] At OpenAI’s scale, that translates into three concrete requirements: Global reach for more than 900 million weekly active users Surely the number refers to the total users of ChatGPT overall, and the fraction of those who use voice features is considerably smaller, is it not? That’s the kind of thing that influences business decisions like knowing how much hardware and software optimization to throw at a problem.
- stuartmemo 5mo agoYeah, that's why they've used "reach" - the total number of users who could be exposed to the feature regardless of engagement.
- deleted 5mo ago[deleted]
- janalsncm 5mo agoTo defend them a little: voice is a little rough around the edges now, so there’s a chicken and egg problem of whether to prioritize improving voice if usage isn’t high partially because it’s clunky.
- notfromhere 5mo agoid rather use the thinking models so the voice mode isnt' useful, i do use voice-to-text more and more just to speed things up though
- Aeroi 5mo agoif anyone is looking to get into this. pipecat is a great open-source repo and community. https://github.com/pipecat-ai/pipecat https://github.com/pipecat-ai/pipecat
- BoxedEmpathy 5mo agoI've been looking at this! Great project.
- pncnmnp 5mo agoI wish I had known about Pipecat a lot sooner. I found out about it a few weeks back, and since Gemma 4 launched, I've been building my own entirely local voice assistant using Gemma 4 + Kokoro TTS + Whisper from scratch - https://github.com/pncnmnp/strawberry https://github.com/pncnmnp/strawberry. Pipecat's smart turn model is really good for VAD - https://huggingface.co/pipecat-ai/smart-turn-v3 https://huggingface.co/pipecat-ai/smart-turn-v3
- AnthOlei 5mo agoWhat do you have going on the hardware side? I want to plug this into hass but don’t know what hardware I need for reasonable latency
- Sean-Der 5mo agoCheck out [0]. You can do 'Voice AI' on small/cheap hardware. It's the most fun you can have in the space ATM :) It's been a while, but posted a demo here [1] [0] https://github.com/pipecat-ai/pipecat-esp32 https://github.com/pipecat-ai/pipecat-esp32 [1] https://www.youtube.com/watch?v=6f0sUEUuruw https://www.youtube.com/watch?v=6f0sUEUuruw
- AnthOlei 5mo agobeautiful demo - is it running fully locally or talking to 3rd party API’s? That box was jaw dropping small
- Dorrell 5mo ago[flagged]
- doctorpangloss 5mo agowhat i learned from making a webrtc+kubernetes game streaming product: - openai is wrong. almost of the issues they described are issues with libwebrtc, not with webrtc, kubernetes, network architecture, etc. the clue was when they said "the conventional one-port-per-session WebRTC model." - there are no alternatives worth trying. everything else open source in the ecosystem, like pion, coturn, stunner, are too immature. - libwebrtc is the only game in town. - they haven't discovered libwebrtc feature flags or how it works with candidates, which directly fix a bunch of latency issues they are discovering. a correct feature flag can instantly reduce latency for free, compared to pay for twilio network traversal style solutions - 99% of low latency voice END USERS will be in a network situation that can eliminate relays, transceivers, etc. it is totally first class on kubernetes. but you have to know something :) this is the first time i'm experiencing gell mann amnesia with openai! look those guys are brilliant, but there is hardly anyone in the world who is doing this stuff correctly.
- jiggawatts 5mo agoSomething I noticed is that companies that are vibe-coding their products miss out on the intelligence that (still) only humans can bring to bear. Just the knowledge cutoff alone puts AI at a serious disadvantage in any rapidly changing field.
- fragmede 5mo agoGPT 5.5's knowledge cutoff is August 2025. Which aspect of WebRTC has meaningfully changed since then?
- mschuster91 5mo agoThe problem is the sheer amount of knowledge out there. Particularly when using niche technologies (which webrtc and web audio still is, when measured by how many people develop using it), it is not surprising that AI doesn't have everything available in its responses, unless you specifically ask it about something you already know it should know.
- 5mo ago
- flakiness 5mo agoShould I or shouldn't I be glad to see zero mention on Codex.
- mock-possum 5mo agoShouldn’t, I think - advanced voice is a surprisingly slick feature, and if you’re someone who feels that they can think and speak more naturally than when they think and type, AI voice transcription is kind of huge.
- gyanchawdhary 5mo ago100% .. as a product designer/developer, i use it heavily for early feature ideation .. i’ll do a loose, exploratory back and forth on a long walk .. then pass the transcript to claude to validate and turn into a spec ..
- furyofantares 5mo ago> Global reach for more than 900 million weekly active users lol, definitely didn't need to know there's 900M weekly users for this post. I mean yeah, there's a lot of users and they serve globally, that's relevant. But this is just pulling out your biggest stat because you can. How many voice users you have would actually be relevant and interesting but, to baselessly speculate on motivation here, might be a number that doesn't add as much fuel to an upcoming IPO as reminded people that you're almost at a billion users does.
- didibus 5mo agoI wouldn't mind waiting longer for answers that would go through a better model with more thinking. As long as it has good support for interrupting and also it doesn't start answering as soon as I pause for 1 second and it's smart about knowing I'm done speaking.
- charisma123 5mo agoIf a transceiver crashes during a stream, how is the active session recovered? Does the system automatically re-establish the context in a new WebRTC session?
- Sean-Der 5mo agoIt doesn't today, but you could with sometime like this [0]. You can save/suspend all WebRTC state and bring it back with the next process. [0] https://github.com/pion/webrtc-zero-downtime-restart https://github.com/pion/webrtc-zero-downtime-restart
- legohead 5mo agoThe low latency is more of a pain point than a good thing, the way they have it implemented. Trying to have a casual conversation with it, as humans we naturally pause, and GPT will take this as you are "done" and start blabbing away. I also suffer from finding the appropriate word I want as I've gotten older and slower, and this fast-voice-gpt just ends up frustrating me more than helping. I have to sit there and think out the whole sentence in my head before I say anything -- not very natural.
- zamadatix 5mo agoI think these are 2 different layers of "latency". The latency in the article is referring to the transport of the audio stream itself while the latency in your scenario is about how quickly to start responding inside the audio stream.
- ericmcer 5mo agoI think he’s saying they are doing an insane level of complexity to shave ~100ms off response times in a scenario where that isn’t important and might even be a negative
- deleted 5mo ago[deleted]
- deleted 5mo ago[deleted]
- zamadatix 5mo agoWhen GP mentioned reducing conversational latency as a negative that made sense (and should probably be done IMO), it just wasn't the same category of latency the article talks about reducing. I.e. increasing "network latency" just makes the conversation feel more and more out of sync, it doesn't change the rate at which the AI will interrupt ("turn latency") because the latter is based on the duration of the pause in the audio stream, not the duration it took to deliver that audio stream. If you meant there is a case where reducing the network latency at the same delivery reliability for a given audio stream is actually a negative then I'd love to hear more about it as I'm a network guy always in search of an excuse for latency :D.
- Sean-Der 5mo agoVery grateful that OpenAI published the article/publicized their usage of Pion[0] a library I work on. If you aren't familiar with WebRTC it's a super fun space. I work on a book WebRTC for the Curious [1] that details how it works. [0] https://github.com/pion/webrtc https://github.com/pion/webrtc [1] https://webrtcforthecurious.com https://webrtcforthecurious.com
- thatxliner 5mo agoslightly unrelated but what’s with storing the entire codebase in the root directory instead of a nested src folder? It makes getting to the README a lot more difficult
- nemothekid 5mo agoThats the default for go projects. Go imports are repository strings (e.g.): import ("github.com/go-sql-driver/mysql") so it's standard to have the library files in the root directory.
- a456463 5mo agoThis is valid criticism. Go fanbois don't like listening to any go criticism. They were all like who needs templates in go. and now go has templates. To me go code looks like somebody vomitted stuff in the root dir and i have to wade through that every time. No namespacing. nothing
- deleted 5mo ago[deleted]
- junon 5mo agoI don't like go as a personal preference but reducing them to "fanboys" is a bit reductive. I'm sure the same could be said about your own favorite language.
- altmanaltman 5mo ago
- testing_auth 5mo ago[dead]
- CrzyLngPwd 5mo agoIt's bad enough having to speed-read the waffle of its written answers; even when told to be concise, the thought of having to listen to it waffle on in its smarmy, sycohpantic fashion makes me want to reach for the sick bag.
- deleted 5mo ago[deleted]
- qrush 5mo agoAm I reading this right that OpenAI is not using Livekit for WebRTC/audio anymore?
- fidotron 5mo agoIt does appear that way. The LiveKit server is not what you would want for this architecture anyway (as they basically say with the SFU discussion), although it does have a lot of useful stuff in the client SDKs.
- geekin23 5mo agothey are not and haven’t been from what I hear since last Jan… I also have some friends work in listed companies on their website within real time divisions but haven’t used livekit, only signed up. It’s kinda shady tbh
- rvz 5mo agoOpenAI uses Go for the networking implementation for the relays and the services, which makes a ton of sense, instead of something immature as TypeScript / Node or whatever. Yet another reason to not consider anything else like that for low-latency networking. Golang (or even Rust and C++) is unmatched for this use-case.
- deleted 5mo ago[deleted]
- nvarsj 5mo agoCan golang do zero copy networking nowadays? In the past golang was terrible at this kind of thing due to allocations and copies of all relayed data.
- bananamogul 5mo ago"something immature as TypeScript / Node or whatever" Node.js's initial release was May 27, 2009 Golang 's initial release was November 10, 2009 They're different, yes, but it's not like
- mghackerlady 5mo agookay, sure, but one is by microsoft, the other by a 25 year old, and another by rob pike. The one by rob pike is going to be infinitely more mature and thought out than a hacky type system on JS because it isn't his first rodeo
- testing_auth 5mo ago[dead]
- jonahs197 5mo agoWho cares? Their company is dying.
- DumpoLumbo 5mo ago[dead]
- DumpoLumbo 5mo ago[flagged]
- logickkk1 5mo agoIMO this probably isn't just about latency. keeping people in voice gives them training data text never will. is that why they were fine going transceiver over sfu and mostly ignoring multi-party?
- tom1IIIl1iIL 5mo agoI think it's better to join some kind of club if you want to make friends?
- Lucasoato 5mo agoWait a minute... I’m genuinely happy that they are sharing this, but keep in mind that realtime audio model from OpenAI are still stuck with the 4o family in terms of capabilities, sadly. I still find them so useful, such a pity that there’s no real competitor in this segment, having the experience a real conversation has helped me so much in expressing ideas and concepts. Still, it’s worth to keep in mind that these are not frontier models, differently from when they were released. (Please Sam, if you read this, release the new realtime audio models)
- dharma1 5mo agoYes the voice part of OpenAI realtime/voice mode is great but it’s pretty dumb compared to newer models and often gets stuck repeating itself. Google’s Gemini flash live 3.1 is better, especially used via the API - it can do tool calling (including to other, even smarter LLMs if you set it up yourself), you can set the reasoning level (even high is still close enough to realtime) and it can ground answers in google search. I love bidirectional voice and right now it’s probably the best option. You can try it in AI studio
- Lucasoato 5mo agoThanks, I’ll try it, even if my experience wasn’t that great with Google models lately (503s)
- dharma1 5mo agoGive it a shot, 3.1 live one in AI studio/API and max out reasoning - not the one in Gemini app it’s an older model. Another option is to use pipecat with their VAD and separate STT and TTS and any (fast) LLM of your choice - but it’s more plumbing and not a true speech to speech model
- stavros 5mo agoHaha, wow, I never thought I'd see a voice model that was too quick, but 3.1 live felt like it responded unnaturally quickly! I'm kind of blown away, I'd want to insert a 100ms delay to make it sound more natural, wow. I never thought I'd see that.
- hnav 5mo agoRFC 9297 support can't come quick enough in browsers. Would obviate having to deal with WebRTC in a client-server scenario.
- mt_ 5mo ago[dead]
- charcircuit 5mo ago[dead]
- devopsengine 5mo agoInspired
- deferredgrant 5mo ago[flagged]
- amirathi 5mo agoI find OpenAI's speech-to-text model the best of the lot. It can handle my & my 5-year old daughter's Indian accent pretty well. I wonder if they run the STT model's output through the current model (that we're chatting with) as a final pass - since the text seem to be well aligned to the current conversation context. For long prompts, I often speak to OAI web/app and copy-paste the text to Claude / Gemini :)
- tracyhenry 5mo agoAfter all these, I still feel their voice AI interrupts quite a lot, especially when I pause just for 0.5 sec. Interestingly, when I tell it to interrupt less, it seems to be better.
- deleted 5mo ago[deleted]
- hiroakiaizawa 5mo agoInteresting. What are the main latency bottlenecks in practice?
- whateveracct 5mo agowhy is the "How" included here? it is often removed
- vjay15 5mo agoThis is such a good write up, WebRTC is one of the coolest things ever! It's kinda genius to use the VIP approach, SFU is also pretty scalable but now they dont even have to do that
- NikolaosC 5mo ago[dead]
- Ozzie-D 5mo ago[flagged]
- zerop 5mo agoI have used voice mode on chatgpt, Gemini, Grok as I use it while driving. Best is from openAI. Natural conversation, smarter and meaningful replies.
- shevy-java 5mo agoI don't like AI in general, and on youtube there are soooooo many horrible videos with voice AI. Having said that, I did notice AI has actually worked for some hobbyist-maintained games for the most part. Example: BG2EE (Baldur's Gate 2 Enhanced Edition). Yes, this is a forgotten game; and I actually have background music as audio rather than listen to the dialogue, save for testing it, but for the most part it worked here. So for poor-ressource hobbyists, AI is actually not totally useless. For youtube I find only horribly crap examples. I don't watch any AI-involved videos (if I can spot it; so much fake on youtube these days, Google does not realise how AI is killing many old users and visitors here).
- maxglute 5mo ago>feels natural if conversation moves at the speed of speech As someone use to podcast at 3x speech and sapi text to speech at much higher rate, listening to AI at human speech is a chore.
- Saline9515 5mo agoI never use the voice mode in the phone app, it's stuck with 4o for some reason. Same with Claude, that uses Haiku. Why can't they use a better model with thinking disabled?
- ath3nd 5mo ago[dead]
- SandeepJawahar 5mo ago[flagged]
- soleiman 5mo ago[flagged]
- Daniel_Visovsky 5mo ago[dead]