4 ms·
Instead of one optimization profile for a whole movie, Netflix is detecting when the shot of a movie changes (the camera "cuts" from one shot to another) and be
by 2bitencryption 6y ago
Instead of one optimization profile for a whole movie, Netflix is detecting when the shot of a movie changes (the camera "cuts" from one shot to another) and beginning a new optimization profile specifically for the content of that shot.
At least, that is my understanding.
- purerandomness 6y agoThat sounds so intuitive that I want to follow up with the question: Why isn't this the default? Is detecting a cut particularly difficult?
- dexterdog 6y agoEven if you have to pay somebody to tag the scene changes manually that cost would be pretty insignificant to Netflix.
- MrMorden 6y agoNetflix should be able to get the EDL, which would probably take care of most of the work.
- Groxx 6y agoCoarsely detecting cuts is relatively simple - look for large frame-to-frame differences (e.g. encode however -> find large frames surrounded by smaller ones -> done, it's as accurate as your perceptual compression is). There are a number of ffmpeg-using tools out there doing this and other "cut to / from black" detection and it's pretty good. Not good enough for a human to say "yeah, these are all scenes", but probably good enough for picking things to re-encode like this. The harder part is the significantly increased compute use due to re-encoding things multiple times, to detect these cuts and to try to find the best encoding. Heuristics there can be arbitrarily complex and re-calculate any number of times. I imagine it hasn't been done earlier just due to cost, though maybe they've recently achieved a better heuristic. edit: ah, great, they link to a "dynamic optimizer" post that goes into this in some detail: https://netflixtechblog.com/dynamic-optimizer-a-perceptual-video-encoding-optimization-framework-e19f1e3a277f https://netflixtechblog.com/dynamic-optimizer-a-perceptual-v...
- myself248 6y agoSeems to me that if the frame-to-frame difference isn't big enough to detect that way, it's not likely to benefit from a new I-frame, yeah?
- Groxx 6y agoThat's the basic idea, yeah. It falls apart in a couple places, e.g. when the cut or fast-fade goes to a very cheap frame like a mostly solid color, and it may not detect stuff like whip-cuts (since a whole chunk of frames are expensive), but so many scenes in so much of media has single-frame cuts that it's well within that "good enough" range. And for dynamic encoding like this: when it's wrong, it's not visually worse in that scene than choosing that sub-par encoding for the entire movie, which has the same "choose the best encoding" problem as individual chunks have. I assume it'd be relatively rare for it to result in anything worse than a one-shot strategy. --- ffmpeg will let you easily do frame-to-frame-diff logic that lets you chop videos into scenes, for example: https://video.stackexchange.com/a/30701 https://video.stackexchange.com/a/30701 I'm not sure how much it handles compressed-frame differences, but it shouldn't be too hard to build around it. Just might be a bit beyond bash-friendly.
- hinkley 6y agoImagine if the compression for the movie “Hero” couldn’t exploit the color themes in each act of the movie. That would be a might bigger.
- nhoughto 6y agoInteresting that it makes sense now because the economics of it have shifted, as the shift away from broadcast and towards individual stream has happened.
- iwasakabukiman 6y agoTraditionally, you just encode the video with a new frame every X frames, regardless of shots. It might sound trivial, but it is still extra computational work to figure out when the shot changes.
- duhi88 6y ago"Variable bitrate" encoding is a thing, which allows chunks of a file to be encoded at a higher or lower bitrate depending on how much is happening. I don't know exactly how it happens, but I assume that each chunk is determined by duration or size, while Netflix's new method determines each chunk by the content. In order for this to be the default, you'd need either humans, or pattern recognition algorithms to identify the chunks. You also need to quantify how by how much each chunk needs its encoding parameters tweaked. Monetary costs aside, that's increasing complexity of your pipeline with relatively small gains. I'd bet that more companies will start looking at similar approaches now that 4K HDR (and 8K) are becoming more common. Probably not worth the R&D for 1080p, but we'll see tricks like this start to trickle down, I'm sure. It's highly unlikely we'll see anything similar in FOSS tools in the near future.
- ehsankia 6y agoRight, these gains are mostly relevant when you have a relatively small library compared to how many views each item gets. For example this would probably not be worth it for Youtube, expect maybe on individual hyper popular videos. Every kb that Netflix can shave off of a popular movie means terabytes of bandwidth saved.