4 ms·
It does a different thing. RTSP is mostly about control and framing. It doesn't specify any particular algorithm to estimate the network's capacity, how a vide
by keithwinstein 7y ago
It does a different thing.
RTSP is mostly about control and framing. It doesn't specify any particular algorithm to estimate the network's capacity, how a video encoder should try to match that estimated capacity, or how to recover from lost packets.
H.265 is a format for compressed video that defines a bit-exact decoder. It doesn't specify the way that an encoder encodes anything into the compressed format, how the encoder should try to match an externally supplied target frame size / bitrate target (while also meeting a latency target), or what the API should be.
The Salsify techniques could work fine with RTSP, and could work fine using H.265 as the coded video format. The special thing about Salsify is really about where the control lies.
Traditionally (in Skype, FaceTime, or the WebRTC.org codebase), there is a drop-in codec with its own control loop (making frame-by-frame decisions), and a congestion-control protocol with its own control loop (making packet-by-packet decisions), and these control loops are at close enough timescales that they end up doing poorly together. And the API to the codec is generally too limited (especially when it's a very general API, as in WebRTC.org, that tries to abstract across pre-existing implementations of VP9/H.264/H.265 to give the application agility across different formats) to achieve the kind of rapid adaptation to network flakiness that you need over these cellular or bad Wi-Fi networks.
Salsify basically says, "hey, if your codec supports a functional-style API, and your transport protocol too, and you can extract the long-lived control state from each of those individual modules and just have one control loop that jointly controls both the codec module and the transport module, you can do a heck of a lot better."
- dillonmckay 7y agoThank you for clarifying that for me.