4 ms·
Low-latency packet video can work incredibly well over a dependable network connection (with a known constant throughput and no jitter), low end-to-end per-pack
by keithwinstein 7y ago
Low-latency packet video can work incredibly well over a dependable network connection (with a known constant throughput and no jitter), low end-to-end per-packet latency, and good isolation between everybody's microphone and speaker. This was mostly solved in the 1990s.
A lot of what makes Skype/Facetime/WebRTC/Chrome suck are the compromises and complexity inherent in trying to do the best you can do for when these things don't hold -- and sometimes, those techniques end up adding latency even when you do have a great network connection.
Receiver-side dejitter buffers add latency. Sender-side pacing and congestion control adds latency. In-network queueing (when the sender sends more than the network can accommodate, and packets wait in line at a bottleneck) adds latency. Waiting for retransmissions adds latency. Low frame rates add latency. Encoders that can't accurately hit a target frame size on an individual frame basis add latency. Networks that decrease their available throughput (either because another flow is now competing for the same bottleneck, or the bottleneck link capacity itself deteriorated) cause previously sustainable bitrates to start building up in-network queues, add latency.
And automatic echo cancellation can make audio incomprehensible, no matter how good the compression is (but the alternative is feedback, or making you use a telephone handset).
Another problem is that the systems in place are just incredibly complex. The WebRTC.org codebase (used in Chrome and elsewhere) is something like a half million lines of code, plus another half million of vendored third-party dependencies. The WebRTC.org rate controller (the thing that tries to tune the video encoder to match the network capacity) is very complicated and stateful and has a bunch of special cases and is written in a really general way that makes it hard to reason about.
And the fact that the video encoder and the network transport protocol are usually implemented separately, by separate entities (and the encoder is designed as a plug-in component to serve many masters, of which low-latency video is only one, and often baked into hardware), and each has its own control loop running at similar timescales also makes things suck. Things would work better if the encoder and transport protocol were purpose-designed for each other and maybe with a richer interface between them (I'm not talking about changing the compressed video format itself; just the encoder implementation), BUT, then you probably wouldn't have access to such a competitive market of pluggable H.264 encoders you could slot in to your videoconferencing program, and it wouldn't be so easy for you to swap out H.264 for H.265 or AV1 when those come along. And if you care about the encoder being power-efficient (and implemented in hardware), making your own better encoder isn't easy, even for an already-specified compression format.
Our research group has some results on trying to do this better (and also simpler) in a principled way, and we have a pretty good demo video: https://snr.stanford.edu/salsify https://snr.stanford.edu/salsify . But there's a lot of practical/business reasons why you're using WebRTC or FaceTime and not this.
- algesten 7y agoWell put! It's also worth mentioning that because of the network conditions, UDP is a better choice than TCP. The reason is that the retransmissions in TCP are not helping with real-time video. However every protocol doing this must have fallback from UDP to TCP because there are a surprising amount of corporate firewalls out there that arbitrarily limits the use of UDP. We worked extensively with webrtc, but are recently switching to our own protocol. The main reason is that webrtc is very complex and generic. You can make a better user experience if you adapt the protocol to target your use case. Another reason is that we can more easily detect firewalls and give the user feedback about it.
- tracer4201 7y agoI bookmarked your link and the paper it links to for reading on my Monday commute. In college, I had one Networks course which discussed IPs, subnets, cidr, UDP, TCP, various higher level protocols, packet headers, etc. There were a couple projects - I think I recall implementing TCP or sliding window. Anyway, that’s the extent of my background knowledge. Where would one start if they want to dive deeper in video transmission or related topics that provide more than the very basic understanding I have?
- hardwaresofton 7y agoThanks for your work on the Salsify paper -- the paper was a great read and the novel approach is pretty great was illuminating. I wonder if it would be possible to get Salsify into the browser using something like web-udp[0]. I don't think the lower level access to the encoder information is available to coordinate network usage... For those who are interested in code, check out the encoder on Github[1]. [0]: https://github.com/osofour/web-udp https://github.com/osofour/web-udp [1]: https://github.com/excamera/alfalfa https://github.com/excamera/alfalfa
- ignoramous 7y agoKeith, thanks for your inputs. Noob question: Where do you think Google's BBR https://ai.google/research/pubs/pub45646 https://ai.google/research/pubs/pub45646 and CMU's HFSC https://www.cs.cmu.edu/~hzhang/HFSC/main.html https://www.cs.cmu.edu/~hzhang/HFSC/main.html fall short? -- For folks unaware, Keith created https://mosh.org https://mosh.org and is an expert in Computer Networks. Here's a relevant talk by Keith on Sprout, a new transport protocol for live video on noisy cellular networks that uses probabilistic inference to predict congestion; and Remy, a program that generates transport protocols on-the-fly in response to network conditions: https://youtu.be/UsCOVF0vDe8 https://youtu.be/UsCOVF0vDe8 And here's a talk at Usenix on Salsify: https://youtu.be/LPj2ffe7Isk https://youtu.be/LPj2ffe7Isk news.yc discussion: https://news.ycombinator.com/item?id=16964112 https://news.ycombinator.com/item?id=16964112 UoCambridge's David McKay on Information Theory, Pattern Recognition, and Neural Networks : https://videolectures.net/course_information_theory_pattern_recognition/ https://videolectures.net/course_information_theory_pattern_...