4 ms·
Seems like even though people went after this, there's not been much innovation, last few decades. Most of the new stuff looks quite incremental. It almost seem
by khitchdee 12y ago
Seems like even though people went after this, there's not been much innovation, last few decades. Most of the new stuff looks quite incremental. It almost seems like we are losing our edge in our ability to build from the ground up. This is because we are part of a system that maintains an eagle's eye on new ideas. We need to take some of that pressure off us so that we can think a bit outside the box.
- gwern 12y agoI wonder how much of this is that relatively simple techniques work very well on small amounts of data, and that more complex techniques don't perform sufficiently better to justify baking them into a compression format? That is, to some extent I subscribe to the compression-as-intelligence school of thought: compressors are tiny little AIs which try to predict regularities in the bitstreams they are given ( http://prize.hutter1.net/ http://prize.hutter1.net/ http://mattmahoney.net/dc/dce.html http://mattmahoney.net/dc/dce.html http://www.danburfoot.net/research.html http://www.danburfoot.net/research.html ). But when we look at the state of the art like ZPAQ, the AI techniques used don't seem to be much more complex than, say, a one or two layer neural network which might as well be from the 1970s. You don't see anything fancy like deep networks or other modern staples like random forests. So this makes me wonder: maybe compression performance has stagnated because we're not willing to provide compression algorithms extremely large amounts of data or runtime, and so simple algorithms really do perform best with the minimal resources we're willing to use for compression. (People are happy to run neural networks on thousands of GPUs with many gigabytes of data and wait weeks for training to finish; can you imagine a compression utility which required that?)
- khitchdee 12y agoI think it's also partly due to the way that the typical dev team develops their algos. They all start with a base version and later push out revisions to their earlier base work. The tendency is therefore to just refine and not go for more game-changing answers. If they executed over longer terms and were more isolated from their markets, we might have seen more range in the changes made.
- vtuulos 12y agoIt naturally follows from the compression-as-intelligence school of thought that building general-purpose compression/intelligence is hard. I very much believe in domain-specific intelligence, and correspondingly domain-specific compression. Here's a practical business use case for lossless compression (in-memory analytics), which I have been developing: http://tuulos.github.io/pydata-2014/ http://tuulos.github.io/pydata-2014/ In contrast to general-purpose encoders, this approach is extremely data-intensive, compressing terabytes per chunk.
- gwern 12y ago> that building general-purpose compression/intelligence is hard. Yes, but my point is. that we seem to be doing better at general-purpose intelligence than at compression despite the apparent equivalence of progress.
- brazzy 12y agoCan you provide any concrete indication that anyone was actually discouraged from researching general-purpose lossless compression by a "system that maintains an eagle's eye on new ideas"? Looks to me like forcing a social explanation on a technical problem. To me it simply looks like there just isn't much possibility of improvement in that field. Ultimately, you're limited by the pigeonhole principle. Special-purpose lossy compression, i.e. video - that's where you see leaps and bounds in active research and improvement.
- khitchdee 12y agoMost of the past algos that have been developed have been done by commercial development teams. This means they are typically driven by the current requirements of the user market. The techniques that get developed therefore tend to be mostly increments to older techniques and since corporates value their IP, these get patented, everyone is happy. If we could step out of this corporate context, we would likely do better on the algos. On the topic of video, the latest stuff is still based on decades old core technology. Its still block-based motion compensated transforms which has been around since at least MPEG-1. Don't see any leaps and bounds in what's used now. No wavelets, no fractals, no other new stuff. Just better optimized versions of the base.
- yason 12y agoThis is no surprise because reading the article, the history of data compression seems to be riddled with patent lawsuits and general blocking of "competitors" to use the same or a similar algorithm.
- derf_ 12y agoWell, part of this is that some of these are just "solved problems". For example, once you have arithmetic coding, you never have to worry about the actual bitwise encoding again (anything on top of that is just speed optimization). That reduces the problem to modeling and probability estimation. There's an inherent limit at the entropy of the source, and usually it is not too hard to get within a few percent of that entropy (as near as we are able to model it). After that, you can do increasingly complex things, but for rapidly diminishing returns. For example, there are definitely lossless audio compressors that can make things smaller than FLAC, but usually only by a percent or two, and they are an order of magnitude or more slower. To me, a lot of the interesting research at this point is exploring the compression/speed trade-off. That's what makes LZMA interesting (it dominates bzip2 on both axes). Google has also done some interesting work in this space. This matters when your goal is to save network transmission time: if you're only sending the data once, anything you do has to be faster than just sending more data on the wire.
- khitchdee 12y agoI don't think it's as simple as you've put it, in terms of solved problems and entropy limits. Part of the quest is to find a transformed space in which the probabilities bin out bit better.
- jbb555 12y agoThe reason there has not been much innovation is that even ZIP files are pretty close to the best you can do for general purpose compression. You might get a few more percent by improving the algorithms but there simply isn't much more compression to be done...