7 ms·
How does the competition stop a highly-specific compressor that is tuned to the particular text file (and not much else) from winning?
by codeulike 6y ago
How does the competition stop a highly-specific compressor that is tuned to the particular text file (and not much else) from winning?
- rockdoe 6y agoIt doesn't. But the "particular text file" is a huge chunk of Wikipedia, so whatever you come up with isn't going to be as "highly specific" as you'd think.
- codeulike 6y agoI think whats confusing me is that they think AI is the solution, so that would tend towards a huge very clever compressor program (think something like GPT3). But presumably the competition also has limits on the size of the compressor - because you want to avoid trivial solutions like embedding the 1GB file in the compressor. So how do they manage this contradiction between wanting the compressor to be very clever while also limiting its size?
- jeffffff 6y agothe size of the program is counted as part of the compressed size
- rockdoe 6y agoJust read the rules. The size of the decompressor is included in the total size of the compressed data.
- codeulike 6y agoRight, but then they're not going to get a very clever compressor then. The ideal AI compressor would encompass all human knowledge. It just seems like they've hit upon a good idea - AI/compression crossover - then shot themselves in the foot by excluding anything genuinely intelligent. I suppose the problem is how to tune the rules to allow AI but not cheating.
- chriswarbo 6y ago> very clever > ideal > all human knowledge > genuinely intelligent > allow AI but not cheating Those all all vague, hand-wavey concepts; open to disagreement, and in some cases might turn out not to exist or make sense. Entire research careers have been spent trying to even define these terms, let alone implement them. Focusing on compression of a particular chunk of Wikipedia eliminates all of that, and gives us a precisely measurable quantity. Is it a perfect defininition of intelligence? No; it was never meant to be. Is it a runnable, measurable, comparable experiment? Yes.
- rockdoe 6y agoThe ideal AI compressor would encompass all human knowledge. That's what they're doing (in the constraints of defining all human knowledge ~ English Wikipedia). More refined models of representing all of our knowledge will beat weaker ones in this test. Whether this is AI or intelligence or not depends on how you define those terms - but to win this competition your compressor has to be able to anticipate the rest of the human knowledge fairly accurately based on having seen part of it.
- codeulike 6y agoOk I see the idea now. The winning program is currently 15meg or so. I can see how that means _something_. But I'm wondering if starting with a 1tb knowledge file might lead to more interesting ai outcomes as it would allow for larger models to be in play.
- chriswarbo 6y agoThe compressor is the thing being measured; there is no separate input file or anything. If you embed 1GB of text in that compressor, then your 'compressed size' will be 1GB (plus whatever else is in the compressor).
- peeters 6y ago> The total size of the compressed file and decompressor (as a Win32 or Linux executable) must not be larger than 99% of the previous prize winning entry. Emphasis mine. You can make a highly specific compressor, but it'll need a highly specific decompressor, which will count against your score. In essence the competition is "create the smallest executable which produces enwiki9 as its output".
- beaconstudios 6y agowhich is equivalent to approaching the kolmogorov complexity of the text, which is deeply interesting research for information theory.
- codeulike 6y agoAha yes, I see
- yreg 6y agoOr on the other side, if there was a procedural generator that would generate the required output from some short seed, that would basically win the game for good. But I guess the possibility of such a generator existing (within the size constraints) for this specific text is pretty much zero.
- goldsteinq 6y agoYou can't produce information out of thin air, so basically your "generator" would need to store all information in the file somehow, and it boils down to writing good compressor to store it.
- yreg 6y agoYou can't produce any arbitrary information out of thin air, but you could produce a specific one by chance. E.g. a Minecraft world seed carries very little data compared to the world which the world generator consistently generates out of it. The catch is that only a miniscule fraction of all the possible worlds can be actually generated since there is a finite amount of coresponding seeds.
- goldsteinq 6y agoStoring arbitrary information via RNG configuration seems improbable.