3 ms·
I expected general-purpose compressors to do better, but I was wrong. For the whole document: 174355 pg11.txt 60907 pg11.txt.gz-9 58590 pg11.txt.zstd-9
by pronoiac 2y ago
I expected general-purpose compressors to do better, but I was wrong. For the whole document:
174355 pg11.txt
60907 pg11.txt.gz-9
58590 pg11.txt.zstd-9
54164 pg11.txt.xz-9
25360 [from blog post]
- EgoIncarnate 2y agoWhich model did you use?
- Intralexical 2y agoIDK why, but BZIP2 seems to do somewhat better that other compression algorithms for natural language text: $ curl https://www.gutenberg.org/cache/epub/11/pg11.txt | bzip2 --best | wc 246 1183 48925 Also, ZSTD goes all the way up to `--ultra -22` plus `--long=31` (4GB window— Irrelevant here since the file fits in the default 8MB anyway).
- pastage 2y agoYou can use a preshared dictionary with bzip2 and zstd so you can get that down alot by using different dictionaries depending on certain rules. I dont know if it helps with literature but I had great success in sending databases with free text like that. In the end it was easier to just use one dictionary for everything and just skip the rules.
- immibis 2y agoAnd here's the opposite - using gzip as a (merely adequate) "large" language model: https://aclanthology.org/2023.findings-acl.426.pdf https://aclanthology.org/2023.findings-acl.426.pdf By the way, if you want to see how well gzip actually models language, take any gzipped file, flip a few bits, and unzip it. If it gives you a checksum error, ignore that. You might have to unzip in a streaming way so that it can't tell the checksum is wrong until it's already printed the wrong data that you want to see.
- anonu 2y agoShouldn't you factor in the size of the compression tools needed?
- Legend2440 2y agoLLMs blow away traditional compressors because they are very good predictors, and prediction and compression are the same operation. You can convert any predictor into a lossless compressor by feeding the output probabilities into an entropy coding algorithm. LLMs can get compression ratios as high as 95% (0.4 bits per character) on english text. https://arxiv.org/html/2404.09937v1 https://arxiv.org/html/2404.09937v1
- Der_Einzige 2y agoThe 1-1 correspondence between prediction and compression is one of the most counterintuitive and fascinating things in all of AI to me.
- mjburgess 2y agoIts only equivalent for a very narrow sense of 'prediction' , namely modelling conditional probability distributions over known data. There's no sense, for example, in which deriving a prediction about the nature of reality from a novel scientific theory is 'compression' eg., suppose we didn't know a planet existed, and we looked at orbital data. There's no sense in which compressing that data would indicate another planet existed. It's a great source of confusion that people think AI/ML systems are 'predicting' novel distributions of observations (science), vs., novel observations of the same distribution (statistics). It should be more obvious that the latter is just compression, since it's just taking a known distribution of data and replacing it with a derivative optimal value. Science predicts novel distributions based on theories, ie., it says the world is other than we previously supposed.
- palmtree3000 2y agoSure it is! If we were trying to compress an archive of orbital data, one way to do it would be "initial positions + periodic error correction". If you have the new planet, your errors will be smaller and can be represented in less space at the same precision.
- 2y ago