11 ms·
I think the mistake here is to require lossless compression. Humans and LLMs only do lossy compression. I think lossy compression might be more critical to int
by slashdev 3y ago
I think the mistake here is to require lossless compression.
Humans and LLMs only do lossy compression. I think lossy compression might be more critical to intelligence. The ability to forget, change your synapses or weights, is crucial to being able to adapt to change.
- version_five 3y agoYeah it makes no sense to say it's inspired by intelligence and then require lossless which is definitionally rote work and not intelligent.
- whimsicalism 3y agoNot true, a smart model could be really good at lossy compression and then you only have to store a small delta to make it lossless.
- ClassyJacket 3y agoI'm no mathematician but I don't believe this is true. Lossless information encoding requires all the original information to be present.
- AnotherGoodName 3y agoArithmetic coding allows you to make a prediction and only provide bits for correction. Have the de-compressor predict the next data based on the outcome so far (a statistical prediction of next data will be lossy as it won't always be correct). If the prediction is correct you need to spend very little to confirm that. If it's incorrect you'll need to spend data to correct it. Arithmetic coding is the best way to make this work. It's also been used by all winning entries of the Hutter prize so far.
- ClassyJacket 3y agoAlright I guess I was wrong.
- slashdev 3y agoTechnically you’re correct. If you can rebuild the original losslessly then all the original information is present, just in a different configuration.
- glitchc 3y agoOr at least reproducible. It could still be compressed.
- vladf 3y agoWhat
- AnotherGoodName 3y agoThat's literally arithmetic coding which is used by all winning entries in the above so far.
- whimsicalism 3y agoyep!
- kaba0 3y agoWhich is the exact same function of the difference between any competitor, so you can actually fairly compare them on the lossy part.. like, people have been thinking about this problem much more than the 3 minutes HNers in this comment section.
- sytelus 3y agoHumans can do lossy or lossless. There are plenty of people who can recite entire Bible or Koran flawlessly.
- Supply5411 3y agoAnd there are humans that can jump 8ft in the air. Doesn't mean it's correct to say that "humans can jump 8ft in the air." Very few people are regurgitating verbatim information.
- kadoban 3y agoThat's true, but it seems unlikely that that's a particularly important part of intelligence. The vast majority of people do _not_ do that type of memorization, are they still intelligent?
- anonylizard 3y agoMany can recite the Koran flawlessly, its short and heavily encouraged in education through rote repetition. Much, much fewer can recite the bible, its many times longer. LLMs can also recite the bible and Koran flawlessly, given how frequent the text appears in their training material.
- RugnirViking 3y ago> LLMs can also recite the bible Is this something you've seen demonstrated? I have no doubt they would be able to recite a lot of sections of it, but there are many sections that are less often quoted, and its a long frickin book. If its just a hunch you have I might give it a go and compare if thats okay I think it would be interesting to interpret what given LLM knows and might miss
- _0ffh 3y agoNow you've given me the idea of coaxing new bible verses out of an LLM. Maybe we can even get one to confabulate a complete additional book! =)
- 3y ago
- mik1998 3y agoLossy text compression has little utility.
- JumpCrisscross 3y ago> Lossy text compression has little utility You're describing every book you've ever read and learned from.
- mik1998 3y agoNo, not at all. I don't read a book to regurgitate the text in an imperfect form later. I read it to learn about the ideas and thought. The syntax itself (ie. text) is not that important.
- _jal 3y ago> The syntax itself (ie. text) is not that important So, you're preferentially discarding information you consider extraneous to your application, distilling it to a smaller representation that retains what you consider important about it.
- JumpCrisscross 3y agoMost image compression smooths skies. Because most people don’t want a regurgitated sky. If you’re an astronomer, on the other hand, you’ll compress away the trees. These are all lossy retrieval (and, inherent to compression, transformation) functions.
- dekhn 3y agoAh, see but what they really meant to say was "Lossy text compression has a little utility" but due to lossy compression the meaning changed to be the opposite, thus proving the author's point.
- barrysteve 3y agoGood books aren't lossy.
- lucb1e 3y agoFeel free to start your own competition where people are awarded money for a system that correctly answers some set of questions, divided by the amount of storage that system needs. Sounds like a tough competition to run objectively, not that that makes it less worth doing, but I can see why the parameters were chosen as they were in 2006.
- AnotherGoodName 3y agoSome background before i explain why your suggestion is "not even wrong". If you can predict the next X bits given previous Y bits of observation you don't need to store the next X bits, the decompressor can just use the same prediction algorithm you just used and write out its predictions without correction. This is the same as general purpose AI at a high level. If you can predict the next X bits given previous Y bits of observation you can make reasoned a decision ("Choose option A now so that the next X bits have a certain outcome"). The above is actually the premise of the competition and the reason it exists. What i've said above is in academic papers in detail by Marcus Hutter et al. who runs this competition. That is lossless compression can be a scorecard of prediction which is the same as AGI. Now saying "they should just make it lossy" misses the point. Do you know how to turn lossy data into lossless? You store some data everytime you're wrong on your prediction. ie. Store data when you have loss to make it lossy. This is arithmetic coding in a nutshell. You can turn any lossy data into lossless data with arithmetic coding. This will require more data the more loss you have. The lossless requirement gives us a scorecard of how well the lossy prediction worked. If you ask for this to be lossy compression you throw out that scorecard and bring in an entire amount of subjectivity to this which is unwanted.
- version_five 3y ago[flagged]
- joshxyz 3y agoYour comment is far more worse, lol. His comment actually supports gp's suggestion, just in a different angle. It looks like LLM's are like compression algorithms with strengths and weaknesses in different things. Losslessness doesnt always equate to usefulness. But yea, maybe a different competition for this.
- AnotherGoodName 3y agoSorry but i hope my long-winded explanation explains it. In computer science we have a way to score how good lossy data is. That way is to make it lossless and look at how much data the arithmetic coder needed to correct it from lossy to lossless. This is a mathematically perfect way to judge (you can't correct lossy data any more efficiently than an arithmetic coder). All the entries here do in fact make probabilistic predictions on the data and they do all use arithmetic coding. So the suggestion misses a key point of CS involved here. I don't mean to be rude about it but the idea does need correcting.
- thelastparadise 3y agoGood lossy compression can be used to achieve lossless compression. (information theory) The more accurate the lossy compression is, the smaller the difference between the actual data (lossless) and the approximation. The smaller the difference, the fewer bits required to restore the original data. So a naive approach is use the LLM to approximate the text (this would need to be deterministic --zero temp with a preset seed), then subtract that from the original data. Store this diff then add back to restore the original data bit-for-bit.
- Xcelerate 3y agoWho downvoted you? This is correct; the better the lossy compression, the better the lossless compression as well.
- CamperBob2 3y agoThat's true until the lossy compression alters the data in a way that actually makes it harder to represent. As an example, a bitstream that 'corrects' a decompressed .MP3 file to match the original raw waveform data would be almost as difficult to compress as the raw file itself would be. It wouldn't be a matter of simply computing and compressing deltas at each sample, because frequency-domain compression moves the bits around.
- hinkley 3y agoIn psychoacoustics, one of the most important aspects of how lossy compression works is that they throw away sounds that a human can't even hear, because it's too short or subtle to be noticed. This is the difference between data and knowledge. Someone with perfect pitch and a good memory knows exactly what Don't Stop Me Now by Queen sounds like, and it's a lot smaller than digitizing an LP record in mint condition. This person can cover that song, and everyone will be happy, because they have reproduced everything that makes that song what it is. If anything we might be disappointed because it's a verbatim reproduction (for some reason we prefer cover songs to introduce their own flavor). If I ask you to play Don't Stop Me Now and you sound like alcoholic Karaoke, you haven't satisfied the request. You've lost. Actually we've all lost, please stop making that sound, and never do that again.
- modeless 3y agoThe problem with lossy compression in a competition is that in order to compare different approaches to lossy compression you have to define a metric for the quality of different lossy reconstructions of the data. This is practically impossible to do in an unbiased way. As soon as you define a quality metric then the competition becomes cheating the metric instead of useful compression. So how do you define a metric that can't be cheated? You add an arithmetic coder after your lossy compressor, which turns it into a lossless compressor, and the size of the losslessly compressed data is your metric for the quality of your lossy compressor. It's the only metric that definitively can't be cheated.
- jharohit 3y agoTed Chiang actually explores this concept in his article on ChatGPT and other LLMs https://www.newyorker.com/tech/annals-of-technology/chatgpt-is-a-blurry-jpeg-of-the-web https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
- hnfong 3y agoSee a related benchmark: http://www.mattmahoney.net/dc/text.html http://www.mattmahoney.net/dc/text.html In particular, Fabrice Bellard is leading with a transformer model for “enwik9” that which is 10 times larger than the one used in Hutter. It’s doing quite well with enwik8 too, but perhaps the issue is that the “economy of scale” with on the fly training hasn’t caught up in the smaller benchmark yet. Transformers are definitely lossy, but afaict Bellard’s entry uses the probabilities generated by the model to create an encoding that used fewer bits
- vintermann 3y agoLossy compression is lossless over the information you decided to care about. You may decide that the sound that's outside the human range of hearing isn't interesting. Or that the colours in your camera sensor that's just thermal noise can be safely ignored. Then you can throw it away and do lossy compression. What you care about, an algorithm can't answer for you. Sometimes you may want to keep the inaudible sounds in a recording to better understand the audible ones. Sometimes you may be glad you kept the noise in a picture as it allowed you to later identify the model of camera. There's no context-free answer to what matters. There's no subject-free answer to what matters, there's always a "who" it matters to.
- mrcgnc 3y agoTrue, but then I would argue (as a layman) that AGI would need to be able to guess what matters in a given context to really be AGI.
- vintermann 3y agoMatters to who? To me? In that case, it sometimes does, sometimes doesn't, depending on how lazy I've been in constraining to the algorithm what I might want. To itself? Then who can say if it isn't doing that already?
- Geee 3y agoThe compression should be semantically lossless, but lexicographically lossy. It should find ways to say the same thing in less characters. I.e "Sky is blue. Ocean is blue." -> "Sky and ocean are blue."
- YetAnotherNick 3y agoLossless text compression is basically something like arithmetic coding[1] on top of probability. In fact the loss metric that LLM directly optimize for is bits required for lossless compression. Loss plotted in papers is generally nits/token. The only reason why LLM are not in the leaderboard because it is not a task for compressing human knowledge. It's a task for compressing wikipedia which is tiny in comparison, and the model size plays a big role. [1]: https://en.wikipedia.org/wiki/Arithmetic_coding https://en.wikipedia.org/wiki/Arithmetic_coding
- mjan22640 3y agomath is a lossless compression