6 ms·
OpenAI’s WebRTC problem
- hnav 5mo agoExactly what I thought when I read the original article, though to be fair WebTransport is barely now entering the mainstream with Safari shipping support this year.
- giancarlostoro 5mo agoProbably because WebTransport is the lesser known alternative to WebRTC.
- est 5mo agoWebTransport requires some speicific server setup. cldouflare doesn't support WebTransport well.
- r2vcap 5mo agoThis is frustratingly one-sided writing. Yeah, WebRTC has limitations, but relying on a standard buys you a lot of correctness and reduces long-term engineering cost. The fact that WebRTC is complicated does not mean it is wrong; it means real-time media over the public internet is complicated. Also, networking is inherently stateful. NAT traversal, jitter buffers, congestion control, packet loss, codec state, encryption, and session routing do not disappear because you put audio over TCP or WebSocket. Pretending otherwise is not architectural clarity. It is just moving the complexity somewhere less visible.
- Waterluvian 5mo ago“How hard can it be?” the strawman asked. It’s 2026 and teleconferencing is still such a shit show. There’s billions of dollars to be had and Zoom is at best mediocre, and it can be as bad as Microsoft Whatchamacallit. I’ve never not seen teleconferencing be a ham handed mess.
- fragmede 5mo agoFacetime does alright in the consumer segment.
- bschwindHN 5mo agoThe most frustrating thing about FaceTime is it sometimes appears to significantly duck audio in order to avoid echoes. I can't predict on which devices it will happen, but it often does when I call my parents and it absolutely destroys the conversation. If they're telling me something and I make the slightest "uhuh" acknowledgment sound, their mic input gets effectively muted for a second or so and I miss what they say.
- jayfeather 5mo agoI’ve found that if the recipient is wearing headphones/earphones, you can freely interrupt them without getting ducked (and vice-versa). Doesn’t help much when you’re calling multiple people / have to be on speaker, but makes it predictable at least.
- bschwindHN 5mo agoWell sure, that's the easy case for echo cancellation (remove the source of echoes entirely). But we've had solutions to this for decades and it doesn't involve ducking the recipients audio. Apple and their billions of dollars should have had this completely solved by now.
- charcircuit 5mo agoQUIC is also a standard.
- tekacs 5mo agoYou might have noticed that the author started the blog post explaining themselves: Like 6 years ago I wrote a WebRTC SFU at Twitch. Originally we used Pion (Go) just like OpenAI, but forked after benchmarking revealed that it was too slow. I ended up rewriting every protocol, because of course I did! Just a year ago, I was at Discord and I rewrote the WebRTC SFU in Rust. Because of course I did! You’re probably noticing a trend. Fun Fact: WebRTC consists of ~45 RFCs dating back to the early 2000s. And some de-facto standards that are technically drafts (ex. TWCC, REMB). Not a fun fact when you have to implement them all. You should consider me a Certified WebRTC Expert. Which is why I never, never want to use WebRTC again. I think that they've done more than enough of 'trying the normal way' to be warranted in having an opinion the other way, don't you think?
- sam1r 5mo agoYes,agreed. I also found it apparently obvious that they have proven their experts worth on this subject matter. Many times, over and over.
- BonoboIO 5mo agoBut ChatGPT said …
- forgotusername6 5mo agoRight but they also state they have never implemented TURN which IMO is a marker of WebRTC expertness. (I haven't btw, just the WebRTC experts I know absolutely have written or worked on at some point a TURN implementation)
- K0nserv 5mo agoIt's not that strange. TURN has two main use cases: peer-to-peer when no viable direct path can be found and working around very strict firewalls. Based on the author's experience the first isn't relevant and the second isn't much of a concern for Twitch and Discord. For the latter case HTTP/3 is helping make TURN unnecessary because you can, as the author observes, run UDP over port 443.
- danans 5mo ago> This is frustratingly one-sided writing Tangential, but by being that, it's also refreshingly human writing, vs the both-sidesy bullet listed AI pablum that's all around us these days. I have zero take on the subject matter, but I like that the article had a detectably human flair. And if it was AI written, god help us.
- awkii 5mo agoThis poor soul. There are few protocols I hate implementing more than WebRTC. Getting a simple client going means you need to quickly acclimate to SDP, TURN/STUN, ice-candidates, offers, peer-to-peer protocols, and the complex handshake that is implemented from scratch each time. I can't imagine re-writing the whole trenchcoat of protocols and unintended "best-practices".
- jgalt212 5mo agoHave you attempted to use the Microsoft Graph API to interact with email?
- edoceo 5mo agoUgh. Who's decided to Graph all the things.
- tempaccount5050 5mo agoIt's way better than the old powershell modules imo. What don't you like?
- Sean-Der 5mo agoWhat platforms were you targeting that you found it painful! Sorry it was frustrating. I hope it’s getting better with education/more libraries. It’s also amazing how easy Codex etc… can burn through it now
- moomoo11 5mo agoi like livekit for this reason and their ceo is cool
- moffkalast 5mo agoThe first time I was able to get a working webrtc datachannel setup with aiortc was when LLMs became a thing, before that it it was pretty much impossible full stop. Nobody knows what or how, there are no examples. It's a horrible protocol that just needs to die.
- Giefo6ah 5mo agoYet another victim of IPv4, and you still find countless detractors of IPv6 on every thread where it's mentioned.
- spongebobstoes 5mo agoIPv4 support is necessary, but IPv6 isn't
- whattheheckheck 5mo agoHow would ipv6 handle it
- tardedmeme 5mo agoYou just send packets to the other party's address and they send packets back to yours. Both parties know their address and you don't need a relay in the middle.
- pocksuppet 5mo agoIt's not really relevant in this case since one endpoint is a massive server farm.
- hnav 5mo agoIt is because most of their complexity is in routing packets. With IPv6 you can just have the thing handling the conversation directly addressable by the client. The last 64 bits of a v6 let you have billions of instances in a region.
- ekr____ 5mo agoThis really isn't the case, because people still have firewalls.
- fidotron 5mo ago> WebRTC is designed to degrade and drop my prompt during poor network conditions You want real time that's what you are going to deal with. If you don't want real time and instead imagine everything as STT -> Prompt -> TTS then maybe you shouldn't even be sending audio on the wire at all.
- telman17 5mo agoYep. Maybe there's some additional configuration I'm missing to mitigate the delay but clients don't seem to want to deal with the delay with STT -> Prompt -> TTS. They'll happily suffer occasional quality issues if the conversation feels "real".
- DonHopkins 5mo ago>Yep. Maybe there's some [dropped] issues if the conversation feels "real". Can you repeat that please? It didn't make any sense. This conversation doesn't feel "real".
- cowsandmilk 5mo ago> You want real time Isn’t the point that OpenAI’s use case does not require realtime? When OpenAI responds, it has most of the audio in advance of when the user needs to hear it. It produces audio faster than real time, so a real time protocol is a bad fit.
- Sean-Der 5mo agoThat is not the case. See get-realtime-translate[0 that's doing it as a trickle instead (not turn based). [0] https://developers.openai.com/api/docs/models/gpt-realtime-translate https://developers.openai.com/api/docs/models/gpt-realtime-t...
- kixelated 5mo agoHello Mr Author here. Apologies that my comment replies aren't as funny. Every low-latency application has to decide the user experience trade-off between quality and latency. Congestion causes queuing (aka latency) and to avoid that, something needs to be skipped (lower quality). The WebRTC latency vs. quality knob is fixed. It's great at minimizing latency, but suffers from a lack of flexibility. We still (try to) use WebRTC anyway, because like you implied, browser support has made it one of the only options. Until now of course! WebTransport means you can achieve WebRTC-like behavior via a generic protocol. Choose how long you want to wait before dropping/resetting a stream, instead of that decision being made for you. And yeah my point in the blog is that often the user wants streaming, but not dropping. Obviously you can stream audio input/output without WebRTC. The application should be able to decide when audio packets are lost forever... is it 50ms or 500ms or 5000ms? My argument is that voice AI shouldn't pick the 50ms option.
- lpln3452 5mo agoI haven't really experienced disconnections while using ChatGPT. Gemini is the frustrating part. Simply backgrounding the app (and the web version too) and resuming it causes the response or the conversation with an assigned ID to disappear. Haha.
- Sean-Der 5mo agoI believe Gemini is Websockets? I have the same experience with heavy/custom applications that try to roll their own media stuff. You run into issues around AudioContext and resumption etc... it's a PITA to have to handle all those corner cases :(
- spongebobstoes 5mo agothis misses a few key things but hits on many others webrtc is a bad protocol, without a doubt. I do like websockets as an easy alternative, but you do need to reinvent decent portions of webrtc as a result I like the idea of MoQ but it's not widely used. probably worth experimenting with, especially as video enters the chat > and then a GPU pretends to talk to you via text-to-speech OpenAI is speech-to-speech, there is no TTS in voice mode > It takes a minimum of 8* round trips (RTT) to establish a WebRTC connection signalling can be done long ahead of time, though I don't see this mentioned in the OpenAI blog. I also saw some new webrtc extensions that should reduce setup time further ultimately though, it comes down to > It’s not like LLMs are particularly responsive anyway I expect to see a shift in how S2S models work to be lower latency like the new voice API models that OpenAI announced to be fair, the new models were released the day after this MoQ blog was published
- Terretta 5mo ago> OpenAI is speech-to-speech, there is no TTS in voice mode Which results in the interesting situation where the transcript isn't what was said: Q: Why do the voice transcripts sometimes not match the conversation I had? A: Voice conversations are inherently multimodal, allowing for direct audio exchange between you and the model. As a result, when this audio is transcribed, the transcription might not always align perfectly with the original conversation.
- Sean-Der 5mo agoResponding to some technical points first, but then after that I do see a future that isn't WebRTC. I don't think it matches where WebTransport+WebCodecs etc is going though. > …but as a user, I would much rather wait an extra 200ms for my slow/expensive prompt to be accurate This is the opposite of the feedback I get. Users want instant responses. If you have delay in generating responses/interruptions it kills the magic. You also don't want to send faster than real-time. If the user interrupts the model you just wasted a bunch of bandwidth sending 3 minutes of audio (but only played 10 seconds) > TTS is faster than real-time https://research.nvidia.com/labs/adlr/personaplex/ https://research.nvidia.com/labs/adlr/personaplex/ Voice AI for the latest/aspirational is moving away from what the author describes. It is trickled in/out at 20ms > We really hope the user’s source IP/port never changes, because we broke that functionality. That is supported. When new IP for ufrag comes in its supported > It takes a minimum of 8* round trips (RTT) That's wrong. https://datatracker.ietf.org/doc/draft-hancke-webrtc-sped/ https://datatracker.ietf.org/doc/draft-hancke-webrtc-sped/ > I’d just stream audio over WebSockets You lose stuff like AEC. You also push complexity on clients. The simplicity of WebRTC (createOffer -> setRemoteDescription) is what lets people onboard easily. Lots of developers struggled with Realtime API + web sockets (lots of code and having to do stuff by hand) ---- I think if I had my choice I would pick Offer/Answer model and then doing QUIC instead of DTLS+SCTP. Maybe do RTP over QUIC? I personally don't feel strongly about the protocol itself. I don't know how to ship code to multiple clients (and customers clients) with a much large code footprint.
- sbrother 5mo ago> …but as a user, I would much rather wait an extra 200ms for my slow/expensive prompt to be accurate I disagree with this SO strongly. I find the conversational voice mode to be a game changer because you can actually have an almost normal conversation with it. I'd be thrilled if they could shave off another 50-100ms of latency, and I might stop using it if they added 200ms. If I want deep research I'll use text and carefully compose my prompt; when I'm out and about I want to have a conversation with the Star Trek computer. Interestingly I'm involved with a related effort at a different tech company and when I voiced this opinion it was clear that there was plenty of disagreement. This still surprises me since it seems so obvious to me that conversational fluidity is the number one most important feature.
- coalstartprob 5mo ago[dead]
- sam1r 5mo ago>> ... I say hi to <strike> Scarlett Johansson <strike> Had a nice chuckle.
- keizo 5mo agointeresting read albeit over my head, but i spent half of yesterday comparing Gemini Live (websockets) vs gpt-realtime-2 and while gpt is super good, seemingly more robust. Gemini connects faster.
- jedberg 5mo agoI have a lot of experience in this area (and some patent applications). For Alexa, the device established a connection back to the server and then kept that open, sending basically HTTP2/SPDY/Something like it over the wire after it detected the wake word. This allowed the STT start processing before you finish talking, so there is only a small delay in processing the last few chunks of your utterance. The answer came back over the same connection. In the case of OpenAI, they can't exactly keep a persistent connection open like Alexa does, but they can use HTTP2 from the phone and both iOS and Android will pretty much take care of that connection magically. The author is absolutely right, a real time protocol isn't necessary. It's more important to get all the data. The user won't even notice a delay until you get over 500ms. Especially in the age of mobile phones, where most people are used to their real time human to human communications to have a delay. (If you work at OpenAI or Anthropic, give me a shout, I'm happy to get into more details with you)
- aenis 5mo ago"The author is absolutely right, a real time protocol isn't necessary. It's more important to get all the data. The user won't even notice a delay until you get over 500ms" Not my experience, running around 6,000 conversations per day with voice, with webrtc + cascading (stt/llm/tts) architecture. Maybe I misunderstood your comment, but that 500ms is basically the floor of a stat of the art voice implementation these days - if you are lucky and don't skimp, and do various expensive things like speculative decoding and reasoning. 450ms on the LLM pass alone. Every ms counts in commercial applications of voice ai. If you add 200ms or 300ms to that, it really degrades the conversation. We do a lot of voice stuff to support our business, largely with unsophisticated, non technical users. Last year's attempts, with measured turn to turn latencies of around 1200ms-1500ms, led to a lot of user confusion, interruptions, abandoned conversations and generally very unpleasant experiences. We are at around 700ms turn to turn now, depending on tool usage needed, and its approaching an OK experience, rivalling an interaction with an actual human. We are spending quite a lot to shave another 100ms off that. We do expensive, wasteful things such as speculative LLM passes, we do speculative tool executions (do a few LLM inferences as the user speaks, but don't actually execute non-idempotent tool calls before you know that that LLM pass is usable and the user did not say anything important at the tail end of their sentence) just to shave 100-200ms. When someone says 500ms is irrelevant I am sure they are describing some other use case, not human-to-AI voice interactions. In my experience with voice AI, the problem is not with some occasional dropped webrtc packets. The real hard problem is with strong background noises, echo, and of course accents. WebRTC with its polished AEC implementations helps quite a lot at least with echos. I get the protocol is a major PITA to implement at OpenAI scale, but for anything but hyperscale applications there is lots of good, viable solutions and commercial providers (say, Daily for instance) that make it a no problem. The real problems to solve are still elswhere. But boy, add 500ms to my latency budget and you've killed my application.
- Aeroi 5mo agothere are a lot of extremely smart people that have come back to webRTC time and time again because it continues to solve problems other methods and protocols can't. with saying that, quic is certainly interesting going forward, but i primarily stream voice + vision at 1fps so it just makes sense, and websockets fail and are insecure at scale for this use case (see https://www.daily.co/videosaurus/websockets-and-webrtc/ https://www.daily.co/videosaurus/websockets-and-webrtc/) . also just listen to sean in this thread, dude knows whats up.
- fy20 5mo agoNice fun article. Gives me Why The Lucky Stiff vibes.
- Aeroi 5mo agoI run the gemini live api over a mesh hosted managed webrtc cloud. works fantastic, and Ive been running it for 2 years. you can try websocket, handle ephemeral keys, ect ect. but when you speak with people running voice agents at scale in this space, many of the issues are solved with webRTC and pipecat and the many resources allocated to solved problems in this space. It certainly feels overkill, and it probably is, but once connection is established, it's pretty magical. the startup time and buffering has been solved for quicker voice connections too, https://github.com/pipecat-ai/pipecat-examples/tree/main/instant-voice https://github.com/pipecat-ai/pipecat-examples/tree/main/ins... (video is harder)
- brcmthrowaway 5mo agoThis is interesting. Does niche knowledge in this area command $1mn salary?
- ec109685 5mo ago> “Here’s a million dollars to implement WebRTC for the fourth time” “Hell no” > “Umm…”
- hnav 5mo agoIt can, in general knowing how to shuffle packets according to RFCs is a pretty decent gig. Pretty much every hyperscaler ends up building various LBs and the learning curve is too steep to just toss randos at it unsupervised, but at the same time it's not necessarily inventing anything new most of the time.
- schappim 5mo ago"WebRTC is the problem" is bait; his real claim is "WebRTC has annoying transport-layer characteristics that hurt cloud Voice AI scaling"... Having just had to tackle this again for my own startup, I'm reminded about what you would lose by ditching WebRTC - the audio DSP pipeline, transmit side VAD, echo cancellation, noise suppression, NAT traversal maturity, codec integration, browser ubiquity etc.
- molszanski 5mo agoI remember using webrtc data channel for p2p video. Browser to browser UDP is neat :) fun memories. Thank you for the read
- nutanc 5mo agoMost of the problems happen because we want to simulate human conversations. While thats a good goal to have, another approach is to let the user know clearly they are talking to a bot. You will be surprised at how accomodating users can be when they know they are talking to a bot and want their queries resolved.
- elephantum 5mo agoMy biggest frustration with WebRTC was precisely captured in the article: even if you don't need p2p and your video source is the process on the same host with your browser, you have to dance around connection setup like you're on a different side of a planet
- yalok 5mo agoThere're tons of ways to fine-tune WebRTC that it wouldn't corrupt audio in poor network - it has all of the controls to smoothly trade-off latency vs quality. Not just NACKs - FEC, disable PLC/Acceleration/Deceleration, larger JB (tons of parameters) etc. Most of the glitches I heard with OpenAI's Voice were not WebRTC related - but rather, to my ear, they sounded more like realtime issues with their inference - which is a very different component to optimize.
- mohsen1 5mo agoI've been using LiveKit which is also WebRTC based and it is super annoying when speed slows down or speeds up at times when connection is not robust. We were using OpenAI's websocket based RealTime audio which was way too slow. So I don't know which one is better. Generally our users like the LiveKit implementation better so maybe WebRTC with enough clever hacks is the answer. This blog was super insightful for me to understand what are the root problems in the current implementation though.
- gozzoo 5mo agoI didn't understand - why is WebRTC good for Google Meet and not good for all other conferencing apps?
- solatic 5mo agoWhy does the voice need to be sent to the server? Why not perform speech-to-text on-device? Is the p10 phone/laptop not capable of this yet, despite every "dictation" feature I see in every modern OS?
- omcnoe 5mo agoAn eventual goal is likely to allow interacting with the LLM directly via audio tokens in input/output skipping tts and stt completely.
- splittydev 5mo agoAmazing read. Blog posts rarely keep my attention like this one.
- geetee 5mo agoRefreshing to read something not in the voice of an llm.
- perryizgr8 5mo agoHow is OpenAI Voice mode any different than a Whatsapp call? Ignoring the part that there is a GPU on the other side instead of a human. But what is the technical challenge in the voice call portion? It seems like that has been a solved problem for a long time now.
- vachina 5mo agoWhy worry for OpenAI. Their product will fail if it doesn’t work. Then they will figure it all out later.
- jongjong 5mo agoI've long had the feeling that WebRTC was intentionally over-engineered. Over-engineered and poorly documented. IMO, tech standards should be simple and minimal and people should be able to implement whatever they want on top. I tend to stay away from complex web standards.
- yugoslavia4ever 5mo ago[dead]
- fps-hero 5mo ago> But nope, WebRTC has no buffering and renders based on arrival time. Like seriously, timestamps are just suggestions. It’s even more annoying when video enters the picture. I felt that comment my bones. Why would anyone possibly have the need to know actual presentation timestamp and how that corresponds to actual realtime? Evidently, no one working on WebRTC has had to synchronise data streams from varying sources before with millisecond accuracy. I was doing a demo for a video stabilisation using a webcam and IMU module in the browser. It turns out the latency between video->rtc->browser and sensor->websocket->browser are wildly different and not constant. The obvious solution would be to send UTC timestamps for the sensors data and synchronise in browser. Not possible, the video has no UTC timestamp reference. When you have control of both sides of the WebRTC pipe, you can do fun things like send the UTC timestamp of the start of the stream, but this won’t solve browser jitter. It worked well enough for a POC but the entire solution had to be reengineered.
- toast0 5mo agoAt least the WebRTC library (not sure about browser integration) can do some a/v sync. RTP audio and video both have timestamps; but of course they have different frequencies and epochs. RTCP sender reports include an RTP time and an NTP time, so you can correlate them. Personally, I'm not thrilled with how webrtc modulates playback to try to synchronize the two streams, so the SFU I work with doesn't send NTP timestamps in the sender reports or we just don't send sender reports; I can't recall the details atm. Part of the problem may be that our SFU always send audio immediately, but video gets buffered and paced. For 1:1 calls not using the SFU, a/v sync seems to work and was not controversial when we enabled it.
- Michael666 5mo ago[dead]
- stackedinserter 5mo agoJust give me mpegts in <video> element, I'm dying.
- 0xbadcafebee 5mo agoExcellent writeup. I wish we had awards for blog posts when the person is a domain expert in the post's subject.
- AdityaAnuragi 5mo agoBrowser API reliability in general has a lot of undocumented edge cases — WebRTC isn't alone there.
- thutch76 5mo agoI didn't make it all the way through the post, but I have to say I think he fundamentally understands the purpose of WebRTC. He calls himself an expert, and yeah he's written SFU's in go and rust and different companies ... but his technical credentials do not mean he's correct. Maybe it's a comprehension issue on my end, but he seems to associate things like stun and dtls as related, compounding issues (particularly in round trip time), but they are really orthogonal. Also, he spends too much time talking about how you can't resend packets, and reiterates that point by stating they tried really hard (at discord?). That's where he lost the plot, imo. The RTC in WebRTC is about real time communication. Humans will naturally prefer the auditory experience of an occasional dropped packet, vs backed up audio or audio that plays at an uneven rate. To clarify, I'm talking about human speech here. If you want to tolerate packet loss, use a protocol based on tcp instead of udp. But you know what happens when you send audio over poor network conditions with tcp? There will be pauses on the receiving end as it waits for the next correct packet. Let's say the delay is multiple seconds. What should the receiving end do when packets start flowing again? Plays the clogged audio at a natural clock? Attempt to play the audio back at a higher rate to "catch up" with any other channels? People, humans, do not generally prefer that experience. Forget about WebRTC for a minute, but instead think about tcp vs udp for voice. Voip has been based on udp since the 90's for a reason.
- entrope 5mo agoI think you're not really engaging with his point, which is that RTC is a poor fit for communicating with an AI agent. I didn't read the blog as claiming that WebRTC is bad for what it is, only that it's a (very) poor choice for a voice-to-AI application.
- ricardobeat 5mo agoOnly if you expect to interact with the agent in a turn-taking format, with (possible) pauses between every turn. ChatGPT’s voice mode is like speaking to someone in real time on a voice call, not input -> output.
- thutch76 5mo agoThat's fair. My attention wanted and I lost the plot. However, I don't think having an agent on one side necessarily changes anything. Network problems are not predictable, particularly on mobile, so the human is still very likely to experience a poor auditory experience on a tcp connection.
- gafferongames 5mo agoJust use UDP
- urbandw311er 5mo agoThis feels like an underrated comment. Anyone here got the technical chops to address it?
- singpolyma3 5mo agoIf you're just doing STT and TTS why would you not do that locally and steam text?
- ukanhaupa 5mo agoBecause local STT and TTS is not good enough and LLMs understand it much better?
- OfekSh 5mo ago[dead]
- jazzyjackson 5mo agoOh is this why 1 800 CHAT GPT is trash now? It worked great when I started using it months ago. Last few times I've called the bot constantly interrupts herself, or stops as if I'm interrupting. I can't get a single full sentence out of her so I stopped calling. I've experienced super deranged behavior out of 1800CHATGPT too, when I was just bored and called to ask how she's doing, what's her day like, she spiraled into laughing maniacally. It was unsettling, that was just before the service became unreliable, so I'm really curious what changed about the architecture.
- pedalpete 5mo agoI used to work in WebRTC back in it's earlier days and our team developed the open-source rtc.io. (https://github.com/rtc-io https://github.com/rtc-io) I never would have imagined that OpenAI is sending the full audio of a request to their servers. I had always assumed the audio was transcribed locally and then sent to the server. The only reason I can think they'd want the full audio is for later model training, which, ok, fair-enough, but this can still likely be done without the limitations of WebRTC.
- FlamingMoe 5mo ago[dead]
- dalemhurley 5mo agoI get what you are saying, I honestly thought it was me who didn’t understand.
- dalemhurley 5mo agoWhat would actually be really interesting, text to speech on the device, you could easily stream text to the client which could generate the voice in realtime, far less bandwidth, latency is not really an issue.
- urbandw311er 5mo ago> You speak into the microphone, it gets sent to one of OpenAI’s billion servers, and then a GPU pretends to talk to you via text-to-speech. Neato. People (including this article) keep talking about OpenAI realtime like it’s a STT - LLM - TTS pipeline but I think this is a fundamental misunderstanding of how the model works. My understanding is that it accepts (and outputs) actual raw audio waveforms. Which, for me, is the sheer joy and wonder of the thing.