11 ms·
Video Live Streaming: Notes on RTMP, HLS, and WebRTC
- liuliu 4y agoThere is a new Low Latency HLS protocol: https://developer.apple.com/documentation/http_live_streaming/enabling_low-latency_http_live_streaming_hls https://developer.apple.com/documentation/http_live_streamin... that claims can get down to 300ms latency.
- vr000m 4y agoLow latency HLS is creating partial segments by bucketing 200ms of frames instead of 6s segments in standard HLS. Whereas in webrtc, the endpoint is sending the frame as soon as it is ready. The apples apples comparison here is 0ms (in webrtc, no send side buffering) vs 200ms (in low latency HLS) or 6s (in standard HLS). This is independent of latency of the endpoint from CDN or source. Another distinction is playback wait time, i.e., how quickly upon joining can it start rendering video. I’m assuming the full reference picture (typically, an IFrame or a golden frame depending on the codec) in low latency HLS is only available at the start of each 6s segment and not in partial segments. So joining a live stream, the receiving endpoint would have to wait at most 6s before rendering. Similarly in webrtc, it’s up to the system to generate a reference frame at regular intervals, as low as every second. Or to do it reactively, a receiving endpoint can ask the sender to send a new reference picture. This is done via a Full intra request, the wait time can be as quick as 1.5 times of the round trip time (as new codecs can generate a new iframe instantaneously upon receiving a request). There’s a slight cpu penalty for this which means that the sender getting too many full intra requests may typically throttle the response to 1s. So Apples to Apples comparison for wait time would be up to 1s for webrtc vs 6s for HLS.
- jalino23 4y agoreal world around: HLS: 10-30s LL-HLS: 2-6s WebRTC: 0-2s. add couple more startup seconds depending slow is your signaling
- vr000m 4y agoAgreed. My calculations are totally skipping finding/updating the manifests or exchanging the SDP Offer/Answers.
- giantrobot 4y agoAn HLS segment can carry any number of GOPs. A GOP length of FPS/2 or FPS/4 will get you an I-frame pretty quickly allowing the GOP to be decoded. MPEG-DASH can do the same IIRC. So there doesn't need to be a segment length delay in playback and typically is not.
- liuliu 4y agoYou don't necessarily need to wait for a reference picture to start playback. Modern codecs all support "intra-refresh", which allows you to reconstruct a reference frame from a set of existing frames. With that, you can set periodic intra refresh much lower than 6s keyframe intervals.
- phlhr 4y agothis will be the winner purely because Apple devices are the lowest common denominator. That is, one must support iOS and if iOS only supports LL-HLS but other devices optionally support LL-HLS then LL-HLS is the winner.
- kwindla 4y ago[post author here] In addition to what vr000m said above, I'll just add that when you make HLS chunks smaller, you're reducing the leverage you get from HLS's core design decisions. I tried to cover some of this in the post. One way to think about this intuitively is that HLS and WebRTC are opposite ends of one important trade-offs axis. HLS is about delivering media streams in a way that scales as cost-effectively as possible. WebRTC is about delivering media frames at the lowest possible latency. These are very different goals, and given current infrastructure and standards it's not possible to have your cake and eat it too, here. That may change in the future as low-latency video becomes more and more important. QUIC, for example, is a new approaches to building out a full stack that works around some of the fundamental tradeoffs that exist today. The result is that pushing HLS segments down to 200ms is not at all a clear win. We'll see what happens as HLS implementations improve. And I should say that my brain has been warped by working on UDP/RTP stuff for a long time. But my bet is that using 200ms HLS segments is, for most real-world users, going to make HLS worse in every way than WebRTC would be, for the same use cases. (That's definitely true today with the early implementations of LLHLS.)
- mchusma 4y agoI appreciated your post so thank you. I would be more interested in understanding why you don't see <1s HLS chunk sizes as working in most cases? I feel like p99 real world latency stuff would show some natural buffer sizes.
- kwindla 4y agoThe smaller the chunk (file) sizes, the less benefit you're getting from pushing the chunks through a CDN. More requests will come back to your origin server. And there's a lot of complexity in the pipeline from encoder -> origin server -> CDN that's mostly hidden (which is good) until you hit big performance cliffs (which is bad). This is something that I think most people trying to implement low latency HLS and DASH have struggled with. It's not only the connection to the client that can stall. Your CDN can stall internally, too, waiting on chunks. And, in fact, if your CDN never ever has any internal performance issues, that's probably an indication that it's configured in such a way as to make the costs of delivering video through the CDN pretty much the same as the costs of delivering that same data through cascading WebRTC media servers! Also, the CDN -> client link is TCP, so you're giving up the ability to just drop packets. TCP is going to do its lossless/ordered thing for you, which again is great for most of what we do on the Internet but starts to actively work against you when you're trying to get down to very low latencies. Does that make sense? I tried to cover some of this in the footnotes. Apologies for not doing a better job.
- fiestajetsam 4y agoWebRTC is not the future of low latency live streaming... at least not outside of video conferencing. It's incredibly complex as a specification, has limitations and numerous issues that set limits in how scalable it can be. Conversely, for HTTP segment based formats like HLS and DASH have limits due to their design (and that of HTTP) which set how low latency can actually go. Where's the future? A likely candidate may come out of the "Media over QUIC" work in the IETF, which already has several straw man protocols in real world use (Meta's RUSH, Twitch's WARP). It'll be a few more years before we see a real successor, but whatever it is will likely be able to supersede both WebRTC and HLS/DASH where QUIC and/or WebTransport is available.
- Sean-Der 4y agoI am heavily biased toward WebRTC. Here is my take on it though! > It's incredibly complex as a specification What is complex about it? I can go and read the IETF drafts, webrtcforthecurious.com, https://github.com/adalkiran/webrtc-nuts-and-bolts https://github.com/adalkiran/webrtc-nuts-and-bolts and multiple implementations. QUIC/WebTransport seems simple because it doesn't address all the things WebRTC does. > has limitations and numerous issues that set limits in how scalable it can be https://phenixrts.com/en-us/ https://phenixrts.com/en-us/ does 500k viewers. I don't think anything about WebRTC makes it unscalable. ----- IMO the future is WebRTC. * Diverse users makes the ecosystem rich. WebRTC supports Conferencing, embedded, P2P/NAT Traversal, remote control... Every group of users has the made the ecosystem a little better. * Client code is minimal. For most users they just need to exchange Session Descriptions and they are done. You then have additional APIs if you need to change behaviors. Other streaming protocols expect you to put lots of code client side. If you want to target lots of platforms that is a pretty big burden. * Lots of implementations. C, C++, Python, Go, Typescript * The new thing needs to be substantially better. I don't know what n is, but it isn't enough to just be a little better then WebRTC to replace it.
- sumy23 4y agoHaving attempted to WebRTC as a generic video transport, I can say that WebRTC has insurmountable problems. The two biggest issues are: 1) Lack of client-side buffering. This is a benefit in real-time communication, but it limits your maximum bitrate to your maximum download speed. It’s also incredibly insensitive to network blips. 2) Extremely expensive. To keep bitrate down, video codecs only send key frames every so often. When a new client starts consuming a video stream they need to notify the sender that a new key frame is needed. For a video call, this is fine because the sender is already transcoding their stream so inserting a key frame isn’t a big deal. For a static video, needing to transcode the entire thing in real time with dynamic key frames is expensive and unnecessary.
- aidenn0 4y agoI'm a bit surprised to not see RTP mentioned anywhere. It was a major standard for real-time voice last time I was looking at these things, and it was starting to get some use for videoconferencing.
- vr000m 4y agoA future blogpost will talk about all the networking work that went into optimizing for these large size participation. Webrtc uses a few protocols. RTP is very central to it, ICE and SDP are also very important protocol for NATs/firewalls and capability selections. While we use webrtc as a shortcut to refer to the suite of protocols and didn’t call out RTP specifically for that reason.
- jhugo 4y agoRTP is the fundamental media transport protocol in WebRTC (which is not a protocol, but rather a suite of protocols working together in a defined way). Basically all web videoconferencing uses RTP already.
- aidenn0 4y agoOkay it's funny because I had vaguely remembered something like that, but a "C-f rtp" on the wikipedia page for WebRTC didn't yield any hits, even though SIP (over websockets) was prominently mentioned, so I just figured the similarity in RTC and RTP had caused me to misremember. I assume WebRTC includes STUN/TURN/ICE (negotiated over SIP?) then for traversing NATs? The last time I was really into networking was 2001-ish so that stuff was still around the corner, but I kept up with my reading for a few years after that. I also had some of these acronyms refreshed when setting up Jingle, which uses XMPP instead of SIP, but establishes an RTP connection much like traditional VOIP would use.
- jhugo 4y agoWebRTC doesn't proscribe a signalling mechanism at all. SIP is sometimes used, Jingle (XMPP) sometimes, or sometimes it's just a custom protocol exchanging SDPs (or equivalent structures) over WebSockets or a REST API. WebRTC itself is RTP (DTLS-SRTP), ICE (incl. STUN/TURN), codecs & related parameters, capture mechanisms, all bundled up into a Web API.
- at_a_remove 4y agoThis all looks like a lot of Very New Design Decisions. I tend not to be an early adopter on these, although it is mostly moot because I am no longer in the video delivery business. I have been considering streaming sets of short films, commercials, trailers, and movies to friends in a live stream via a pipeline of comparatively hoary old tech: 1) A playlist of multiple different video files in VLC as a source for ... 2) OBS Studio, which produces RTMP to be consumed by ... 3) nginx, which calls ffmpeg to produce multiple birates and resolutions to be rebundled as HLS to be sent to ... 4) a "live TV" channel of my own specification as an input to JellyFin, which can be read by ... 5) various clients on Roku, Apple TV, Firestick, Chromecast "apps" At this point, I don't think the industry will ever really settle down to something manageable.
- j1elo 4y agoWhy steps 2 and 3? I mean, I don't totally see the reasons for not consuming the video files directly with ffmpeg, instead of going in such roundabout way to bring the media into it.
- deleted 4y ago[deleted]
- at_a_remove 4y agoWell, and understand I haven't yet attempted this, but apparently the discontinuities between switching files has caused a lot of problems downstream for Jellyfin, so OBS Studio would exist as kind of a way to smooth out the transitions.
- wuyishan 4y agoI do something like this by using https://github.com/ffplayout/ffplayout-engine https://github.com/ffplayout/ffplayout-engine
- at_a_remove 4y agoThank you, that looks like that could replace steps one, two, and three with a single step!
- jzer0cool 4y agoIn each of these approaches will any of these (or something else) also provide approaches to getting the stream saved for viewing again later on a remote server? I'm guessing I need some client which also connects to the stream and save which can then be served again. Or is still a solution already baked in available?
- pavlov 4y agoDaily's API will let you record the stream as a MP4 video file stored on Amazon S3 where it's immediately available after the live stream ends. Everything happens automatically on the same server that encodes the live stream. (The original post is written by Daily's CEO.)
- vr000m 4y agoPavlov’s comment is correct. I came to add that soon the stream can be stored on customer’s own S3. Ergo, you’d be able to do a call in real-time, store it on your S3 account and make it available for streaming. On your own S3, this would be a multipart upload.
- rexreed 4y agoIf we use Vimeo for video hosting would there be some way to automatically post to Vimeo or some other way to get the file up there and available for immediate viewing? Other approaches besides S3?
- pavlov 4y agoIf Vimeo offers an ingest API using either RTMP(S) or HLS, that would be one way to get the stream from Daily directly to them without any extra processing step in between.
- dbrueck 4y agoFrom a technical perspective, that's more or less baked into HLS: these are just discrete files being served over HTTP, so you can save the files to disk along with the manifest file and you have the content. In practice, there is a little more to it. For example, when DRM is enabled, you need some way to preserve the decryption keys. And for live content, the manifest file usually just tells the client about a sliding window of files, so you need a tiny bit of additional client side logic to pay attention to this fact. One cool thing about DASH/HLS is that you can do some pretty complex mixing of content - you can build a traditional TV-channel like experience that mixes live and prerecorded content, you can replace and inject ads, you can make live content immediately available for on-demand playback, etc.
- difosfor 4y agoWe're using DASH and experimenting with LL-HLS to get reliable and affordable (using CDNS) end to end latency of 3 seconds with which you can already interact live with your audience without too much trouble. Latencies down to about 1 second are also feasible already at the cost of some client side buffering. So I would have expected to see more information about DASH than HLS here. It's usable on all platforms except iPhones and Apple TV where we're forced to use HLS at higher Latencies for now and hopefully later also DASH or LL-HLS once that's reliably usable. See: https://liveryvideo.com https://liveryvideo.com
- kwindla 4y agoNice! I am a bit skeptical about "down to about 1 second" being achievable with DASH or LL-HLS reliably. Of course, I could be wrong. And a lot depends on the definitions of "about" and "reliably," as well as your user cohort (where they are in the world, etc). :-) The reason I didn't write much about DASH is that the basic concepts are the same for both DASH and HLS. And my sense for the last couple of years has been that most of the momentum in the ecosystem as taking place around HLS/LL-HLS. But I could be wrong about that, too.
- difosfor 4y agoIn some areas of the world where network and/or device performance are limited such a low latency will be impossible to achieve indeed. Inter regional broadcasting also comes with a latency penalty, but with good CDNs like provided by our partner Akamai you can already get a 1 second latency in many cases. That being said the buffers will be very small and we don't recommend using that yet in general. HLS and DASH are similar indeed, but I think the main reason that HLS is still used a lot is that it's the lowest common denominator; you can get it to work everywhere. Perhaps in the future that will be LL-HLS, or something else entirely, but for now most really low latency broadcasting that I'm aware of is using DASH with HLS as a fallback (i.e: CMAF). But I could also be wrong about that of course :-)
- rollcat 4y ago> It's usable on all platforms except iPhones and Apple TV where we're forced to use HLS at higher Latencies for now and hopefully later also DASH or LL-HLS once that's reliably usable. Interesting, since it was Apple who designed & built both the original HLS and the LL extensions. What's currently preventing LL from being usable there?
- dbrueck 4y agoGreat article, thank you! Minor nit about footnote #10 (the very bad network failure mode): it really depends on where the bottleneck is and how the client has been implemented to react - there's a huge spectrum of client implementations out there, ranging from nearly dumb to those having wizard level heuristics. If it's the user's last mile connection (between e.g. their home and their ISP), then the big HLS/DASH/etc buffer translates into a lot of time to react. So clients have the option to shift quite low - and do so quickly - if there are some very low bandwidth options, in theory even switching down to an audio-only or nearly-audio-only stream if one is provided, and can also choose to be optimistic/aggressive to resume playback as soon as one full chunk is downloaded - and some implementations will even resume playback when less than a full chunk is downloaded. The client side logic has a lot of latitude here to balance fast start/resume times vs sustaining playback. When the bottleneck or failure is elsewhere, HLS can be incredibly durable. For extremely high profile events, for example, there are typically multiple CDNs involved, multiple sources going to independent encoders, etc. So an HLS/DASH client might talk to many different servers on a given CDN, as well as servers on alternate CDNs, and even grab what amount to being different copies of the stream spit out by different encoders. It's not uncommon for a client to be testing different CDN endpoints throughout playback to migrate away from congestion automatically.
- kwindla 4y agoThat's a really great point about multi-CDN architectures. Partly because big parts of WebRTC are not standardized (session setup signaling, of course, but also in practice lots of necessary state management) it's a little bit hard to imagine how to build an equivalent for WebRTC. Relatively recently, I would have said that our experience running large-scale WebRTC stuff in production made "core" infrastructure failure relatively low on our list of concerns. The two components of the last mile connection, on the other hand, are always a huge pain point because of the long tail of bad ISPs and bad Wifi setups. However ... many of us who try to deliver always-available video services got something of a wakeup call in November and December last year when AWS had two pretty big outages two months in a row.
- vdnkh 4y agoTwitch latency in low-latency mode is ~2s (you can check for yourself in the settings cog menu -> advanced -> video stats) because of "prefetch segments", which are delivered via HTTP 1.1 streaming body responses. The client requests this segment and long polls while frames are delivered as they're being encoded. It's simpler and cheaper than WebRTC or Apple's low-latency HLS parts, and scales to hundreds of thousands of viewers. The downside is that it's not part of the HLS specification so support for it needs to be bespoke, but it is a proven technology that wasn't covered here (and to my knowledge, the most widely used low latency HTTP solution by # of users). Twitch has been using this for years.
- kwindla 4y agoThe Twitch stuff is really, really good. The reason I didn't talk about it in the post is because the basic ideas that go into reducing the latency of any HLS/DASH implementation are the same. Smaller segment sizes, chunked transport, and smart encoding and prefetch/buffering implementations. But ... while there's lots of terrific engineering underpinning the Twitch approach, there's no way around the real-world limits to that approach. As a result, on good network connections you'll usually get ~2s latency. But not so much on the long tail of network connections. If your use case can gracefully accommodate a range of latencies for the same shared session across different clients, that's fine. If it can't, it's not fine. Plus, I don't think you can get down very much below 2s. (I don't have access to any Twitch/IVS internal data, so I don't know what latencies they see globally across their user base. But I've done a lot of testing of this kind of stuff in general.) And, yeah, Apple. Agreed.
- vr000m 4y agoAgreed, all of the early HTTP based streaming was about HTTP progressive download. Microsoft’s smooth streaming was dominant in the early 2000s used it too. Glad that Twitch has been using it. For me, DASH and HLS are just manifests or playlist definitions, of how best to find the resource. It is not dissimilar to webrtc signalling, i.e., go to a resource to discover where to get other parts of the resource.
- fulafel 4y agoSome history reminders. Internet video conferencing: we used to have multicast and realtime videoconferencing in the 90s: https://ee.lbl.gov/vic/ https://ee.lbl.gov/vic/ Before that we had analog videoconferencing, eg AT&T Picturephone service went GA in 1970.
- vr000m 4y agoUnfortunately modern routers made multicast unrounable, otherwise, we’d all be using mbone
- fulafel 4y agoMbone was tunneled, and expanded to native IP multicast to save bandwidth in networks that supported it. Meanwhile current networks use multicast (non routed, granted) all the time for eg service discovery and lower level stuff like finding the MAC address corresponding to a IP address.