4 ms·
Realistically it would probably be a major refactor to adapt the WebRTC.org codebase to a Salsify-style design. The WebRTC.org codebase is about a million lines
by keithwinstein 7y ago
Realistically it would probably be a major refactor to adapt the WebRTC.org codebase to a Salsify-style design. The WebRTC.org codebase is about a million lines of code (about half of which is vendored third-party code, e.g. libvpx) so any serious refactor is probably out of the realistic capabilities of a university-based research group. For comparison -- the entire Salsify codebase is about 16,000 lines of C++ (including the codec), plus about 7,000 lines of assembly we took from libvpx. It does a heck of a lot less than WebRTC.org, but it's a lot easier to prototype new designs that way.
Salsify is mainly about the benefits you can get if you (a) extract the control loop out of the video codec, and make it expose a functional-style API, (b) use a Sprout-like congestion control algorithm that tries to follow evolving network capacity quickly and estimate "how many bytes can we send right now while trying to maintain a bound on end-to-end delay", and (c) have a single control loop that works every frame and has the choice of whether to send (1) a frame whose coded length is already known and is about equal to what we think the network can handle [it is very hard to get this from any existing video encoder on a single pass!], or (2) no frame at all.
You can certainly do this within the WebRTC protocol, but doing it within the WebRTC.org codebase is going to be a lot of work. :-( I think unfortunately the codebase has some pretty deeply ingrained assumptions about how the control flow is going to go. Inverting that (as we propose), I don't think is an easy incremental change.
Now, you might ask whether there are some incremental improvements you could make to WebRTC.org to get 80% of the benefits of Salsify without a major refactor. E.g., maybe you don't need the functional API if you just make some tweaks to the rate-control algorithm and congestion control. I don't think we know for sure though, but I suspect that for somebody already experienced with that codebase, probably yes, there are gains to be had. There's another question about whether the gains of Salsify (which are on a particular type of flaky cellular network) might come at the expense of costs on other types of (dependable?) networks. To be confident on that we'd probably need to try this stuff much more broadly and on real people (this is the motivation for Puffer).
- microcolonel 7y agoThank you for this informative answer. :- ) I guess a complementary question is in order: since you believe that the WebRTC protocol itself is not incompatible with the mechanisms that enable Salsify's performance, how much work would it be to adapt Salsify incrementally into a WebRTC implementation? (Ignoring mandatory codec support for a moment). I've been developing an Opus codec mode for Bluetooth A2DP, and I have a similar sort of situation. I've been looking (casually) in to using more interesting rate control based on channel performance to improve QoS with Bluetooth A2DP. I think Opus has a property you could only dream of: strict limits on frame size. :- ) Added: just found that the video codec uses the VP8 bitstream, that was not immediately clear to me; seems to me that that is a substantial selling point that should be made more clear on the webpage. Since your interface doesn't require any change to the bitstream of a conventional codec, it's a heck of a lot closer to public use than I initially thought!
- keithwinstein 7y agoThis is an interesting question! I think getting our codebase to use the WebRTC framing and setup would probably not be that hard. We could probably use the same libraries that WebRTC.org is built on (libjingle, etc.), and maybe even just take a lot of their code. I don't think that means that a Salsify sender would be interoperable with, like, a Chrome receiver, though (even though we're just using VP8) -- we'd have to implement at least the receiver side of Salsify's congestion-control protocol inside WebRTC.org/Chrome, and Salsify sometimes likes to encode a VP8 frame that has to be interpreted relative to a certain (prior) decoder state. On #2, honestly I've worked with libopus a bit for Puffer and it's just really pleasant to work with. As you say, a lot of the difficulties of interacting with a video encoder you just don't have in this context. It's also pretty easy to get "gapless/clickless" back-to-back playback of audio excerpts that were encoded completely independently, unlike with video where this is a huge pain in the neck and usually requires a SAP/IDR/closed GOP/sequence header+I-frame+P-frame (which takes a lot of bytes so you can't do it very often). See https://github.com/StanfordSNR/puffer/blob/master/src/opus-encoder/opus-encoder.cc https://github.com/StanfordSNR/puffer/blob/master/src/opus-e... if you are interested for more.