4 ms·
Since the h265 relies heavily on operations which are not easily differentiable, such as translation of patches of images, together with a pretty complicated bi
by vletal 5y ago
Since the h265 relies heavily on operations which are not easily differentiable, such as translation of patches of images, together with a pretty complicated binary format, I'd be pretty amazed if the NN actually learned anything meaningful at all.
- svantana 5y agoInterpolated translation is continuous and easily differentiated. There's lots of work on machine-learned video codecs already, from Nvidia, Qualcomm and others.
- vletal 5y agoIt's not difficult to propagate gradients while translating an image. Learning "pick a 8x8 patch from (145,17) apply X to it and translate it by (4,-8)" from data is on different level, is not it? The premise was using e2e learning to avoid patent issues. I am sure that with some preprocessing you can plug a NN inside the deciding process and learn very meaningful stuff.
- zinekeller 5y ago> There's lots of work on machine-learned video codecs already, from Nvidia, Qualcomm and others. I've tried them, the current state of the art means that they're only useful on relatively static things (some shaking etc) while spike up to AV1-level bitrate to reach perceptual similarity when the movement is too much. Maybe in the future (or whatever under the wraps concoction Nvidia, Qualcomm or another player have), ML-based video codecs will surpass handtuned codecs, but it's not (yet) the present state.