3 ms·
I've been having this "you don't want to use TCP" conversation a lot lately with people who are building real-time voice + LLM applications. Almost everybody wh
by kwindla 2y ago
I've been having this "you don't want to use TCP" conversation a lot lately with people who are building real-time voice + LLM applications. Almost everybody who hacks together a voice + LLM prototype starts with WebSockets. Which I totally understand. WebSockets seem like they should do what you want.
And WebSockets do work fine while you're testing.
But WebSockets are TCP, so when you roll things out to real-world users, latency is higher than it should be a lot of the time. Connections randomly drop. You have to try to figure out why different OS configurations are behaving differently. You start building your own keepalive logic. You start trying to figure out how to build metrics and monitoring for audio-over-WebSockets.
The answer is to use UDP, and use a protocol built on top of UDP designed for real-time media (WebRTC).
It's been interesting to me that I've had this exact same conversation with several dozen engineers over the past few months. It's a good reminder that things that seem obvious when you've been doing something a long time aren't obvious to people new to a domain. Even if they are very experienced in some adjacent domain. (In this case, mobile app and web app development.) It's made me think about where my knowledge boundaries are and what "obvious" things I'm ignorant about. (Lots of them, I'm sure.)
If you're interested in WebSockets, WebRTC, and why UDP is the right low-level approach for real-time media, I wrote a a primer about that a few months ago here:
https://www.daily.co/blog/how-to-talk-to-an-llm-with-your-voice/ https://www.daily.co/blog/how-to-talk-to-an-llm-with-your-vo...