4 ms·
Indeed, that's what I meant by: > Is it so that they can better distribute the packaged encodes to their CDN? The Packager, as far as I can tell, is not actua
by dkarp 5y ago
Indeed, that's what I meant by:
> Is it so that they can better distribute the packaged encodes to their CDN?
The Packager, as far as I can tell, is not actually a cpu intensive process. That article mentions that the bottleneck was caused by IO which was resolved with the Virtual Assembler. It's not doing any encoding at all. It's stitching the encoded video together and packaging it with the audio/subtitles/metadata.
It seems like this could be done at the edge without ever actually packaging the whole file. My guess is the same as yours though, the reason they're actually packaging the encodes is to send them to their CDN. But it's still a guess as maybe the actual reason is that the Virtual Package doesn't work for some other reason.
It could just as well be to allow testing the packaged encode like you would any other deployment artifact before distributing.
- jasode 5y ago>The Packager, as far as I can tell, is not actually a cpu intensive process. [...] It's not doing any encoding at all. A Netflix employee can chime in but I assume cpu intensive since the desired objectives of Packager can only be achieved by transcoding which implies a cpu intensive process. E.g. Netflix wants to use AV1 as one the output formats for Packager. So any process that converts a codec-specific format to other coding-agnostic formats will require cpu. (e.g. Apple ProRes on the input files and AV1 on the output files.) Also more evidence that significant cpu is involved is this: >The overall ProRes video processing speed is increased from 50GB/Hour to 300GB/Hour. From a different perspective, the processing time to movie runtime ratio is reduced from 6:1 to about 1:1. [...] All the cloud packager instances now share a single scheduling queue with optimized compute resource utilization. The 300 gb/hr would be ~83 MB/sec which is way under the disk throughput S3 can provide. AWS says 100 Gbit/sec is possible which would be ~10 GB/sec : https://aws.amazon.com/premiumsupport/knowledge-center/s3-maximum-transfer-speed-ec2/ https://aws.amazon.com/premiumsupport/knowledge-center/s3-ma... >It seems like this could be done at the edge without ever actually packaging the whole file. But on-the-fly-ephemeral "virtual package" instead of a realized output disk file would lead to constant network traffic between the Netflix CDN appliance at the various ISP datacenters back to AWS S3. I guess this is technically possible but not sure what you gain with this alternative architecture instead of the CDN appliance just downloading one "package" file at off-peak hours (4:00am).
- dkarp 5y ago> A Netflix employee can chime in but I assume cpu intensive since the desired objectives of Packager can only be achieved by transcoding I think you're maybe confusing the Packager with the Encoder. Transcoding (encoding) happens before packaging and is distributed. That's what the "encoded chunks" I've been talking about are. There is a different package for each encode, this is in the third linked article from the OP: https://netflixtechblog.com/high-quality-video-encoding-at-scale-d159db052746 https://netflixtechblog.com/high-quality-video-encoding-at-s... The package part seems similar to what is done by MKVMerge (https://en.wikipedia.org/wiki/MKVToolNix https://en.wikipedia.org/wiki/MKVToolNix) + the stitching of the chunked video encodes. There's very little processing necessary compared to encoding. MKVMerge will give you an mkv file from a H.264 video, audio files and subtitles in milliseconds and it doesn't matter how big the files are as it's just a container. They use a different container, but it's the same idea. You wouldn't have to read from S3, you could just push your encoded chunks to your CDN instead and use the edge server to virtual package them.
- jasode 5y ago>I think you're maybe confusing the Packager with the Encoder. Transcoding (encoding) happens before packaging The way I used "transcoding" was to refer to Netflix's Packager process of converting (in their words) "codec-specific" elementary stream format to "codec-agnostic" with extra frame metadata. I should have used a different word than "transcode" to encompass that (especially if the input and output files are the same "codec" but just different containers) ... but whatever the underlying process is, it implies (some) cpu constraints because they're only processing at 83 MB/sec throughput from SSD disks on S3. My laptop doing a simple "mux" type of operation with ffmpeg or MKVMerge can concatenate streams into another container greater than 400 MB/sec. >You wouldn't have to read from S3, you could just push your encoded chunks to your CDN instead and use the edge server to virtual package them. The Netflix blog says the Packager is scanning/analyzing the input file for exact frame start and stop times and storing that knowledge as extra metadata to enable future clients to randomly skip around the video. Just focusing on that one algorithm tells us it's not something we want to do repeatedly (and virtually) on all the edge CDN servers. (I'm reminded of analogous situation in mp3 vbr format that doesn't have exact time frame timestamps built-in for random seeking. Therefore, skipping to exactly 45m17s of a 60 minute mp3 takes a long time as the audio player "scans" the mp3 from the beginning to "count up" to 45m17s. One can build an "index" for fast mp3 random seek but that requires pre-processing the whole mp3 that's more cpu intensive than a simple mux operation.)
- liuliu 5y agoMaybe they simply haven't got there yet. Packaging and deliver to CDN would be technically simpler comparing to maintaining the virtual packager + chunks at CDN level (not to mention the virtual packager if needs to run at CDN level, requires smarter CDN such as Cloudflare Worker).
- fragmede 5y agoThe "edge" is a more recent invention, or rather, the level of abstraction now available to the public for "the edge" lends itself much more readily to such things when designed from scratch. If Netflix were to Greenfield/rebuild it from scratch today, from what I'm reading off Netflix's blog, your proposal seems reasonable. It depends on what their internal abstraction on what their edge actually looks like in practice but I'm guessing it's simultaneously more and less advanced compared to eg Cloudflare's edge workers but institution inertia means "if it ain't broke" is a guiding principal all its own. If you wanted to get a job at Netflix, propose that change, and implement it, I'd bet they'll reward you handsomely for it.