4 ms·
From http://prize.hutter1.net/ http://prize.hutter1.net/: Being able to compress well is closely related to intelligence as explained below. While intelligence
by pitdicker 3y ago
From http://prize.hutter1.net/ http://prize.hutter1.net/:
Being able to compress well is closely related to intelligence as explained below. While intelligence is a slippery concept, file sizes are hard numbers. Wikipedia is an extensive snapshot of Human Knowledge. If you can compress the first 1GB of Wikipedia better than your predecessors, your (de)compressor likely has to be smart(er). The intention of this prize is to encourage development of intelligent compressors/programs as a path to AGI.
The Task:
Losslessly compress the 1GB file enwik9 to less than 114MB. More precisely:
- Create a Linux or Windows compressor comp.exe of size S1 that compresses enwik9 to archive.exe of size S2 such that S:=S1+S2 < L := 114'156'155 (previous record).
- If run, archive.exe produces (without input from other sources) a 10^9 byte file that is identical to enwik9.
- If we can verify your claim, you are eligible for a prize of 500'000€×(1-S/L). Minimum claim is 5'000€ (1% improvement).
- Restrictions: Must run in ≲50 hours using a single CPU core and <10GB RAM and <100GB HDD on our test machine.
- tromp 3y ago> compressor comp.exe of size S1 that compresses enwik9 to archive.exe of size S2 such that S:=S1+S2 < L That criterion is rather more complicated than just taking the size S2 of archive.exe There is no logical meaning to the sum of the size of a compressor and its output that I can see. Disregarding the compressor size would make this contest easier to understand as simply trying to determine the Kolmogorov Complexity (i.e. information content) of enwik9. I looked for the motivation of including compressor size and only found this in the FAQ: > By just measuring L(D)+L(A), one can freely hand-craft large word tables (or other structures) used by C and D, and place them in C and either D or A. By counting both, L(C) and L(D), such tables become 2-3 times more expensive, and hence discourages them. Discouraging word tables seems like a weak and somewhat arbitrary justification for complicating the measure of merit. I don't think the contest would be any less interesting if the nature of the compression would be disregarded. One could even argue that compression of Wikipedia taking way more resources than decompression is justified by having to perform it only once, while its result could be decompressed millions of times.
- viraptor 3y ago> There is no logical meaning to the sum of the size of a compressor and its output that I can see. There is if you have a time limit. Otherwise you could spend a week computing a better result and embed that directly into the compressor - for use only if you're competing that Wikipedia file.
- moritzwarhier 3y agoWhy is a time limit needed? And it doesn't matter if all of Wikipedia is embedded in the compressor — if you can reconstruct the original using only the decompressor and the compressed data, the process can be as expensive as you like. That's the exciting part. It's also a vague connection to what's exciting about LLMs, regarding the amount of information they can reproduce with comparatively small size of their weights. But since decompression must be lossless, it's unclear to me if this approach could really help here. Edit: all of this is already addressed in the FAQ at http://prize.hutter1.net/ http://prize.hutter1.net/ Especially the paragraph "Why lossless compression?" is interesting to read.
- dTal 3y ago>since decompression must be lossless, it's unclear to me if this approach could really help here. All lossy compression can be made lossless by storing a residual. A modern language transformer is well suited to compression because they output a ranked list of probabilities for every token, which is straightforward to feed into an entropy encoder.[0] On the other hand, LLMs expose the flaw in the Hutter prize - Wikipedia is just too damn small, and the ratio of "relevant knowledge" to "incidental syntactic noise" too poor, for the overhead of putting "all human knowledge" in the encoder to be worth it. A very simple language model can achieve very good compression already, predicting the next word almost as well as a human can in absolute terms. The difference between "phone autocomplete" and "plausible AGI" is far into the long tail of text prediction accuracy. Probably a state of the art Hutter prize winning strategy would simply be to train a language model on Wikipedia only, with some carefully optimal selection of parameter count. The resulting model would be useless for anything except compressing Wikipedia though, due to overfit. [0] A practical difficulty is that large language models often struggle with determinism even disregarding the random sampling used for ChatGPT etc, when large numbers of floating point calculations are sharded across a GPU and performed in an arbitrary order. But this can easily be overcome at the expense of efficiency.
- vlovich123 3y agoRegurgitating something verbatim seems less like intelligence vs summarizing the important details. Additionally, a human can even choose the requirements for lossy compression in the first place - do I want a 1 sentence summary of what I read, a 1 paragraph summary, a synopsis, a description of the themes, how the themes relate to other works by the author/creator, etc etc. So I fail to see the fundamental connection between intelligence and lossless compression unless you just choose to define intelligence in some way that connects it to lossless compression.