4 ms·
The thing about video compression is that it's already very compute intensive. Most devices have dedicated hardware cores for encoding and decoding in order to
by xipix 3y ago
The thing about video compression is that it's already very compute intensive. Most devices have dedicated hardware cores for encoding and decoding in order to save power. They work exclusively with integer values, chiefly unsigned 8 and 10 bit values with sequential and block memory transactions.
Humans not only design for compression quality, we also really focus on cost. And we design direct to what is cheap in hardware: integer adds, shifts, binary logic.
I do believe a different paradigm of ML is necessary for this particular application. The neural primitives used need to be more closely related to transistors than to matrix multiplications.
- echelon 3y agoThat's said as if matrix math isn't going to become the most important silicon-enabled process in the world ever. The gravitas will shift to enabling matrix math rather than algorithms being designed for special hardware. Matrices are the new paradigm.
- xipix 3y agoMatrix math units are a smart use of silicon for many ML applications, true. But Von Neumann compute isn't going away. Neither are fixed function modules for shading, ray tracing and massively parallel GPGPU type applications. Also not video coding. My point was that if training could emit an efficient video codec as a network of logic gates, rather than as a dense array of neural network weights, it might just produce something practical for video playback on mobile devices and video processing at hyperscale.
- david-gpu 3y agoSince modern chips from cell phone application processors to datacenter GPUs all have significantly invested real estate to accelerate matrix operations, it only makes sense to take advantage of it, whether it is optimal for the task at hand. That is exactly the same path that led us to GPGPU, which in turn derived to using GPUs to accelerate neural nets. In every step of the way it was about repurposing existing hardware that was designed for something else and was thus suboptimal.
- bradleyjg 3y agoMost videos are only going to be viewed a tiny number of times, often just on the devices they were taken on. It’s probably not worth spending too much effort on compressing these. A small number of videos are going to be viewed millions of times. For these few videos it’s worth the compute to do the best possible compression. There’s no reason the same video formats and decoding hardware can’t be used for both. The high compute version can use an ML model to figure out how to optimize the bandwidth budget across every scene while the cheap version uses some easy heuristics.
- 2h 3y ago> best possible compression This is a common incorrect assumption. You have to balance ratio with speed, otherwise you end up with something that spends more time decompressing than actually downloading, as is literally the case today with XZ. this is why alternatives like Zstandard are getting popular. 9 times faster decompressing for 10% worse ratio.
- bradleyjg 3y agoApples and oranges. The video standards constrain the encode to something that can be decoded in real time by specified decoders.