4 ms·
It's not entirely clear to me what the job of the Packager is here as opposed to a Virtual Package. After chunk encoding and Virtual Assembly with an index fil
by dkarp 5y ago
It's not entirely clear to me what the job of the Packager is here as opposed to a Virtual Package.
After chunk encoding and Virtual Assembly with an index file, are Netflix actually packaging the encoded video/audio/subtitles/metadata into a single file for each encode that is pushed to their CDN? If so, then why is the Packager even necessary? Why not go one step further and create a Virtual Package as well?
Is it so that they can better distribute the packaged encodes to their CDN?
- jasode 5y ago>It's not entirely clear to me what the job of the Packager is here. In the thread article's 1st paragraph is a link to a previous article with more details of Packager: https://netflixtechblog.com/packaging-award-winning-shows-with-award-winning-technology-c1010594ba39 https://netflixtechblog.com/packaging-award-winning-shows-wi...
- dkarp 5y agoRight, and I actually read that. Maybe I should rephrase my question. I'd like to know why is a Packager necessary instead of virtually packaging and presenting an interface of a packaged file? Then you'd be able to keep your encoded chunks as chunks instead of having to download them and then upload them again to MezzFs.
- jasode 5y ago>why is a Packager necessary instead of virtually packaging and presenting an interface of a packaged file? Maybe I'm parsing your question wrong but perhaps the confusion is cleared up if we remember that the "CDN" in the Netflix architecture diagram is not physically stored at AWS S3 : https://miro.medium.com/max/350/1*A5PR2QJ7STPbUd2xTg6z_g.png https://miro.medium.com/max/350/1*A5PR2QJ7STPbUd2xTg6z_g.png See: https://www.google.com/search?q=netflix+cdn+appliance+%22open+connect%22 https://www.google.com/search?q=netflix+cdn+appliance+%22ope... And a recent 2021 article says the Netflix CDN appliance has ~280 terabytes of local disk storage: https://dev.to/gbengelebs/netflix-system-design-how-netflix-onboards-new-content-2dlb https://dev.to/gbengelebs/netflix-system-design-how-netflix-... Thus, the "output" video files of Packager is eventually physically transferred to geographically distributed ISP datacenters. So to attempt to reconstruct Netflix's thought process... Once we work backwards from the need to eventually store copies of video content on distant Netflix CDN appliances, the question becomes which type of video file to store: (a) the codec-specific format (good for archival storage that can generate new downstream formats) -- but not optimized for client playback such as fast random seeking, a/v sync across dynamic changing resolutions, etc) -- or -- (b) the codec-agnostic format -- which is good for client device seeking, etc The option (a) wouldn't make sense since it you'd still need a cpu-intensive process (the "virtual packager to read chunks" in your words instead of The Packager) running at the ISP appliance to present client-device-optimized a/v streams. You'd have wasteful cpu cycles across multiple appliances creating the same "virtual package". So that leaves option (b) ... which means you need a batch process of some type (what Netflix calls Packager) running against AWS S3 storage to create a package for subsequent distribution to all appliances. Therefore, if your proposal of "virtually packaging and presenting an interface of a packaged file" can be reworded as "on-the-fly generated ephemeral a/v frames from the chunks" , their physical topology and need for efficient use of cpu wouldn't make that an optimal architecture.
- dkarp 5y agoIndeed, that's what I meant by: > Is it so that they can better distribute the packaged encodes to their CDN? The Packager, as far as I can tell, is not actually a cpu intensive process. That article mentions that the bottleneck was caused by IO which was resolved with the Virtual Assembler. It's not doing any encoding at all. It's stitching the encoded video together and packaging it with the audio/subtitles/metadata. It seems like this could be done at the edge without ever actually packaging the whole file. My guess is the same as yours though, the reason they're actually packaging the encodes is to send them to their CDN. But it's still a guess as maybe the actual reason is that the Virtual Package doesn't work for some other reason. It could just as well be to allow testing the packaged encode like you would any other deployment artifact before distributing.
- jasode 5y ago>The Packager, as far as I can tell, is not actually a cpu intensive process. [...] It's not doing any encoding at all. A Netflix employee can chime in but I assume cpu intensive since the desired objectives of Packager can only be achieved by transcoding which implies a cpu intensive process. E.g. Netflix wants to use AV1 as one the output formats for Packager. So any process that converts a codec-specific format to other coding-agnostic formats will require cpu. (e.g. Apple ProRes on the input files and AV1 on the output files.) Also more evidence that significant cpu is involved is this: >The overall ProRes video processing speed is increased from 50GB/Hour to 300GB/Hour. From a different perspective, the processing time to movie runtime ratio is reduced from 6:1 to about 1:1. [...] All the cloud packager instances now share a single scheduling queue with optimized compute resource utilization. The 300 gb/hr would be ~83 MB/sec which is way under the disk throughput S3 can provide. AWS says 100 Gbit/sec is possible which would be ~10 GB/sec : https://aws.amazon.com/premiumsupport/knowledge-center/s3-maximum-transfer-speed-ec2/ https://aws.amazon.com/premiumsupport/knowledge-center/s3-ma... >It seems like this could be done at the edge without ever actually packaging the whole file. But on-the-fly-ephemeral "virtual package" instead of a realized output disk file would lead to constant network traffic between the Netflix CDN appliance at the various ISP datacenters back to AWS S3. I guess this is technically possible but not sure what you gain with this alternative architecture instead of the CDN appliance just downloading one "package" file at off-peak hours (4:00am).