2 ms·
Many encoders have a "two pass" mode that saves statistics from a first, fast pass to guide the second pass. Usually these are very simple statistics, used to c
by TD-Linux 9y ago
Many encoders have a "two pass" mode that saves statistics from a first, fast pass to guide the second pass. Usually these are very simple statistics, used to choose a quantizer and frame type in the second pass. Your method feels like a rather extreme version of this, that can guide the whole RDO search with a very large state. And, of course, rarely are first passes done in parallel. It's super exciting to see people looking at this - I think there's a lot of untapped potential. The WebRTC case looks interesting too, though there you need extremely low latencies (1 frame) so I'm curious to find out how you tackle that.
One downside with your approach is that you have to code many keyframes that are eventually thrown away. Keyframes are often expensive to encode because their bitrate is so much higher. Have you considered "synthesizing" fake keyframes somehow, such as doing an especially fast and stupid encode for them? Also, artifacts from keyframes tend to be a lot different from inter predicted frames, so would encoding a keyframe, plus a second frame, then throwing away both yield a quality improvement at the final pass (at the cost of more wasted computation)?
- keithwinstein 9y agoThanks for your kind words! For the WebRTC/real-time case, here is our current draft (https://cs.stanford.edu/~keithw/salsify-paper.pdf https://cs.stanford.edu/~keithw/salsify-paper.pdf). Would be eager for any comments or thoughts or ideas for experiments to run, as we have a few weeks to revise it before the final version is due. Re: having to code many keyframes that are thrown away, in practice we're using vpxenc for the initial pass, and it's just a heck of a lot faster than our own C++ codec (whose benefit is that it can encode and decode a frame relative to a caller-supplied state). We then use our own encoder to re-encode the first frame of each chunk as an interframe (in terms of the exiting state of the previous chunk), and that one encode ends up being slower than encoding six frames with vpxenc (see figure 4). So in a real system, you're absolutely right that this is another optimization opportunity, but I think for our purposes there's a lot of other low-hanging fruit we'd want to tackle first.