3 ms·
It's quite exciting to see progress on a data driven approach to compression. Any compression program encodes a certain amount of information about the correla
by ChrisFoster 10y ago
It's quite exciting to see progress on a data driven approach to compression. Any compression program encodes a certain amount of information about the correlations of the input data in the program itself. It's a big engineering task to determine a simple and computationally efficient scheme which models a given type of correlation.
It seems to me like the data driven approach could greatly outperform hand tuned codecs in terms of compression ratio by using a far more expressive model of the input data. Computational cost and model size is likely to be a lot higher though, unless that's also factored into the optimization problem as a regularization term: if you don't ask for simplicity, you're unlikely to get it!
Lossy codecs like jpeg are optimized to permit the kinds of errors that humans don't find objectionable. However, it's easy to imagine that this is not the right kind of lossyness for some use cases. With a data driven approach, one could imagine optimizing for compression which only looses information irrelevant to a (potentially nonhuman) process consuming the data.