13 ms·
Compression is prediction
- SpyCoder77 2mo agoSomething Ngrok is doing is working, because first they manage to get Sam Rose of samwho.com and now this? At this point I care more about their blog than their products
- Muhammad523 2mo agoI was rushing to post this and then found out somebody had already
- andai 2mo agoYou snoze, you loze! https://www.youtube.com/watch?v=B6u-FPskfAE https://www.youtube.com/watch?v=B6u-FPskfAE
- sheeeeesh 2mo agoGrant Sanderson has an excellent video on the same topic [0]. It's part of a series that is ongoing. [0] Compression is Intelligence Part 1 - https://youtu.be/l6DKRf-fAAM?si=yyLWq8x4sSRkWd98 https://youtu.be/l6DKRf-fAAM?si=yyLWq8x4sSRkWd98
- zahrevsky 2mo agoI wonder if the author of the article knew about the series, or do they both just independently came across this topic to talk about it.
- cyanydeez 2mo agoit was vaguely in my understanding of information & intelligence with compression; it was also brought up in several of the initial trials against AI companies where they discussed how the AI is akin to compression. So they're both sourcing a bit broader zeitgeist.
- deleted 2mo ago[deleted]
- soulofmischief 2mo agoIt's basic information theory, which has been around since the end of WWII. It's a common topic today because some of its subtle insights are becoming increasingly relevant in our current era of AI, as we learn to understand these black boxes.
- epistasis 2mo agoAnybody working in the field will be very familiar with these concepts.
- larodi 2mo ago…for years. Because it is so apparent if you actually try to look at the problem and what is being solved by it. The extraction of features from a corpus, the features significant to certain solution, is always and since day zero - compression. As this is the definition of compression - efficient and potentially lossless feature extraction.
- internetter 2mo agocommon theory. see https://prize.hutter1.net/ https://prize.hutter1.net/
- AnotherGoodName 2mo agoAnd the Hutter Prize for AI which measures how good AI is by measuring how well it compresses data is over 20 years old now just to really drive the point home.
- tulio_ribeiro 2mo agoIn the article, she credits the 2023 DeepMind paper Language Modeling Is Compression as the source. She also links to this post from 2015: https://colah.github.io/posts/2015-09-Visual-Information/ https://colah.github.io/posts/2015-09-Visual-Information/
- samuell 2mo agoThank you for sharing! Small off topic thing: I recommend removing the bit in the URL from ?si=... and forward, unless you want Google to track every user who clicks this link to your share.
- deleted 2mo ago[deleted]
- throwaway_7274 2mo agoThis perspective is a useful source of intuition against the “LLMs can’t have new ideas, they’re just next-token-predictors” style arguments. What if you shift your perspective to thinking of training as optimization over a vast parametrized family of compression algorithms? Well, it suddenly looks a lot more plausible that “new” “ideas” can emerge from that process!
- throwaway_7274 2mo agoIncidentally, the relationship is bidirectional. You can try it out just for fun. zstd is a pretty crappy language model :)
- glial 2mo ago> it suddenly looks a lot more plausible that “new” “ideas” can emerge from that process This is not intuitive to me. It seems like a "new idea" is something that (almost by definition) isn't in the training set. Can you elaborate a bit? Edit: but perhaps a good model could arise from training, which would be a good idea in the sense that parsimonious ideas are good scientific ideas.
- cyanydeez 2mo agoOnce MP3s were invented, I had the idea for the Apple IPOD; but obviously I didn't have a giant manufacturing wing, the ability to make small hard drives, or anything else. I don't think Apple invented the ipod anymore than I invented it; LLMs likely would have also come to the same conclusion about an ipod like device. Original ideas either dont exist or have a functionally irrelevent definition in comparison with inputing tokens to LLMs to get novel ideas out.
- redhed 2mo agoHow I see it, is if the human brain does lossy compression/prediction of the natural world that learns from its "training set" (sensory inputs) and we have been able to come up with new ideas, then it seems like AI would be able to as well.
- jbay808 2mo ago
- deepsun 2mo ago> compressors and LLMs Why only LLMs? All statistical models are compressor. You can say "model" and "compressor" are synonyms. Article does not mention "embeddings" at all, even though it's commonly viewed as a compression method. Also "encoder" part on "auto-encoders".
- hmokiguess 2mo agoI often wonder how would language fare if we didn't have redundancy in abstractions, why do things get different terms, and if there is such a smaller set that contains everything in a lossless way (english-wise)
- andai 2mo agoSee also: Bellard's Lossless Data Compression With Neural Networks https://news.ycombinator.com/item?id=19589848 https://news.ycombinator.com/item?id=19589848 https://news.ycombinator.com/item?id=27244004 https://news.ycombinator.com/item?id=27244004
- adamgordonbell 2mo agoAnd also the LLM version, and LLMZip https://bellard.org/ts_zip/ https://bellard.org/ts_zip/ https://arxiv.org/abs/2306.04050 https://arxiv.org/abs/2306.04050
- speedgoose 2mo agoI tried to reproduce those results, at least in terms of compression ratios, not speed. However I would say that testing on alice29, enwiki8, text8 data is kinda cheating. Alice in Wonderland and Wikipedia are very likely part of the training data of the LLM models used there. So I tried on HN comments from a few days ago, extracted from the text column of the public HN bigquery dataset. Using RWKV v7 0.1B instead of RWKV v4, I get 0.962 bits per byte on alice29, and 1.156 bits per bytes on the HN comments. Still a lot better than 2.826 bits per bytes of xz level 9.
- adamgordonbell 2mo agoOh wow, so it worked pretty well on data it hasn't seen. That expected but cool to reproduce. Have you seen this leaderboard of sorts[1], and this proposal to change hutter prize[2]? I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric because file sizes are very concrete. They are already beating shannon's numbers using a human prediction for compression, from what i can see. https://github.com/hkust-nlp/llm-compression-intelligence https://github.com/hkust-nlp/llm-compression-intelligence https://gwern.net/hutter-prize https://gwern.net/hutter-prize
- krackers 2mo ago
- pjankiewicz 2mo agoI was thinking about the same topic and the conclusion can be wrong. LLMs are compressors, but compressors are not LLMs. Mixing this can let you believe that you can use a compressor to do the same thing as LLMs, which you cannot. Specifically I was thinking about a way to inject knowledge into LLMs training by using statistical properties of text in such a way that you don't have to train the LLM to achieve some level of predictions. There are actually some papers that inject n-grams statistics as a part of the neural network weights.
- Legend2440 2mo ago>Mixing this can let you believe that you can use a compressor to do the same thing as LLMs, which you cannot. You can, actually! Any compressor can be losslessly converted into a generator, and vice versa. Traditional compressors like gzip are of course very simple and can only replicate rough patterns from the input. But they are technically doing the same thing.
- pjankiewicz 2mo agoI agree that technically they are doing the same thing but in practice LLMs are better compressors than PNGs (learned this while I was researching this topic). That was quite surprising to me.
- vrighter 2mo agoActually, it's trivial. I did it for fun once when I was learning about the PPM algorithm. It took about 15 minutes to reverse the whole thing.
- davmre 2mo agoAny compressor actually can be used, trivially, as an autoregressive language model. Given a context (for LLMs, this would include the entire pretraining dataset, plus the prompt), you compress `context + next_token` for every possible next token. The tokens that co-compress best with the existing context are the 'least surprising' continuations. Choose one of them and iterate. You can easily generate text with gzip this way. It won't be very good text, because gzip compression is not as sophisticated as a transformer + SGD, but the principle is the same.
- variadix 2mo agoThis is a lot less surprising when you learn how non-LZ compressors work, that is, by modeling a probability distribution and using those probabilities to encode information in the minimum number of bits required to transmit the data. A less obvious conclusion is that LZ compressors do this to implicitly, the length of each symbol they could emit (literal or match, etc.) can be converted to the probability distribution the LZ compressor induces, since the number of bits to encode the symbol is related to its probability by the information content.
- duskwuff 2mo agoA common design in compressors is to use LZ as a first step, but to then represent the constant data and/or offset-length pairs from LZ using an entropy coder. Deflate (as used in gzip) uses a Huffman coder. LZMA (as used by xz) uses a predictive range coder. Zstandard can use either Huffman or FSE. Some high-speed compressors like LZ4 skip the entropy coding stage entirely at the expense of compression ratio. Bzip2 is an interesting aversion of this pattern - it uses the Burrows-Wheeler transform as a first pass instead of LZ. Unfortunately, this is one of the major reasons why it's so slow.
- sgsjchs 2mo agoThe first LZ-step pretty much directly maps to BPE tokenization in LLMs.
- versteegen 2mo agoIf doesn't correspond cleanly. I can see why you draw the link, because LZ compression will replace words with symbols but BPE is a non-contextual entropy encoding while LZ is contextual and adaptive and that makes it very different. I think BPE actually has more in common with Huffman encoding.
- sethev 2mo agoThis immediately reminded me of the Hutter Prize (http://prize.hutter1.net/ http://prize.hutter1.net/) - a contest that has run since 2005(?) based on the premise that compression is closely related to intelligence.
- ssivark 2mo agoNope; there is a bit more nuance and the distinction is important. Compression is functionally equivalent to prediction when the data distribution is exactly representative of all future problems. The story changes drastically if you want generalization -- because the test distribution could be arbitrarily different, even if it had the same support! Eg: you observe a rare edge case in your training data and (lossy) compression could simply ignore it. But if you wanted generalization in that particular part of the space -- either because an adversary was testing you, or for design freedom where you choose to build in that specific corner -- then you don't just want data compression, but good prediction performance on a test distribution which peaks in that corner. Assuming that the training data distribution is exactly the distribution you will ever care for is implicitly doing a lot of the heavy lifting in the claim that compression = prediction, and I'm peeved at how much this statement is unthinkingly repeated like a manifesto. There is nothing natural about the training data distribution, especially if the data generation process is exploratory while the downstream usage will be exploitative.
- usernametaken29 2mo agoI think Hutter would vehemently disagree with you on that one ;)
- schopra909 2mo ago100% agreed.
- jbs789 2mo agoThat’s interesting. Also sparked the thought that the assumption only holds if the future looks like the present.
- vanviegen 2mo agoIf your compression algrotihm is deep enough (think LLM), it will capture a lot of abstraction, making it compress well even in future cases that differ from the passed but fit the scheme in some other way.
- adamgordonbell 2mo agoSmall world. I just did a podcast on this same topic, but coming at it from a different direction, ie. me and my neighbor trying to beat the hutter prize for compression. Hutter Prize being where you are paid if you can compress wikipedia small enough. LLMs do very well at that, if, big if, you ignore the cost of initial weights. A cool Claude Shannon story: Shannon wanted to measure how much information is actually contained in ordinary English text. His 1948 theory said such a number must exist, but he had no way to calculate it, because the patterns in English reach across dozens of letters and no equation or frequency table captures all of them at once. So instead of calculating it, he ran an experiment on a person. He took a passage from a novel that the subject had not read, and covered it with a card so only the text already guessed was visible. He asked the subject to name the first letter. If the guess was wrong, he asked again, and kept asking until the subject named the correct letter. He wrote down how many guesses it had taken, revealed the letter, and moved the card one position to the right. Then he repeated the process for the next letter, and the next, through the whole passage. What this produced was not a sequence of letters but a sequence of numbers — one number per letter, recording how many guesses that letter required. Most of the numbers were 1, because someone fluent in English, seeing the preceding text, usually names the next letter correctly on the first attempt. Shannon then argued that this sequence of numbers contains exactly as much information as the original passage. Sounds a lot like next token prediction to me. https://corecursive.com/the-hutter-prize/ https://corecursive.com/the-hutter-prize/ http://prize.hutter1.net/ http://prize.hutter1.net/ https://github.com/hkust-nlp/llm-compression-intelligence https://github.com/hkust-nlp/llm-compression-intelligence https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf
- JohnKemeny 2mo ago> ignore the cost of initial weights Well, then Wikipedia itself is a very good compression that only needs the title to perfectly predict the full article.
- adamgordonbell 2mo agoExcept the LLM generalizes its encoding to all english text where as the copy of wikipedia can only 'compress' wikipedia.
- j-pb 2mo agoI always feel like people leave out the third case of the analogy: indexing The article itself has decision trees for the compression explanation, which is also a lookup index. In each case you try to recognise (re)usable structure. Self-indexing succinct data-structures are a good example of the third side of the coin. So it's a trinity: compression, prediction, indexing
- jparishy 2mo agoCool visuals and breakdown. I wrote something in early 2025 about how LLMs seem to be an emergent behavior of lossy compression, but did not have the knowledge or verbiage at the time to get this detailed. In retrospect my writing seems naive and I'm happy to have found this and the Google paper linked inside. To be a fly on the wall in some of the labs, man. Another thought that came from the same post is that, insofar as we see LLMs as human-style intelligence, they're more like stream of consciousness devices. Essentially incessant talking and buying enough time until you get to a usable answer. I think I associate some subset of intelligence with what you don't say, which is impossible with the SOC-style outputs, so this is something I think about a fair bit. What could maybe differentiate current gen models from next gen is the ability to call tools modeled within the layers themselves, not externally. I think as far as I understand it, model trainers expect the model to do this itself in a way we don't understand or control, like a version of the bitter lesson. But I posit we can model many determinate tools as NNs themselves and figure out how to get the internal states of the LLM to make use of them during inference, e.g. calculators, indexes, citations. Just an enthusiast though, so grain of salt and all.
- kailanb 2mo agoUnrelated to the content: I was really pleased to see that this site defaults to the bare minimum for cookie consent. I reflexively clicked "Reject all" only to see that it was already the default, which threw me off.
- baron3dl 2mo agoI stumbled across a connection between LLMs and compression when researching N-dim polytope emergence in neural networks. Toy Models of Superposition (Anthropic, 2022) suggests that gradient descent can independently discover efficient geometric packing arrangements for sparse features. LVQ compression uses regular lattice structures, including some based on 4D lattices. I found this interesting and wonder whether LLMs have a higher density ceiling, since training and inference don't rely on a fixed lattice and can instead learn their own representational geometry.
- woliveirajr 2mo agoThere is Compression done by Prediction by partial matching [0] There is the Kolmogorov Complexity [1], Normalized Information Distance [2] and Normalized compression distance [3] that correlates those. Finally, there's the Pre-Big Bang Informational Compression and the Delayed Release of Antimatter [4] All big {rabbit/black} holes to lose some time, if you have any. [0] https://en.wikipedia.org/wiki/Prediction_by_partial_matching https://en.wikipedia.org/wiki/Prediction_by_partial_matching [1] https://en.wikipedia.org/wiki/Kolmogorov_complexity https://en.wikipedia.org/wiki/Kolmogorov_complexity [2] https://homepages.cwi.nl/~paulv/papers/chapter08.pdf https://homepages.cwi.nl/~paulv/papers/chapter08.pdf [3] https://en.wikipedia.org/wiki/Normalized_compression_distance https://en.wikipedia.org/wiki/Normalized_compression_distanc... [4] https://philarchive.org/rec/GREPBI https://philarchive.org/rec/GREPBI
- vrighter 2mo agoThis is exactly why I think they are one and the same. It's relatively trivial to just plop a (lossy) machine learned markov chain instead of one learned (perfectly) from the data into PPM. With zero changes to the rest of the algorithm.
- brumar 2mo agoI'll add Minimum Description Length to the mix. Under certain definitions and conditions, it equals the Bayesian Information Criterion plus an extra term, which I consider a very interesting result in this "two faces of the same coin" perspective.
- Xmd5a 2mo agohttps://philarchive.org/rec/GRETIO-35 https://philarchive.org/rec/GRETIO-35 https://quantum-journal.org/papers/q-2020-07-20-301/ https://quantum-journal.org/papers/q-2020-07-20-301/
- deleted 2mo ago[deleted]
- farfatched 2mo agoThis is the thesis behind the "Information Theory, Inference, and Learning Algorithms" course that was taught at Cambridge University. > Why unify information theory and machine learning? Because they are two sides of the same coin. In the 1960s, a single field, cybernetics, was populated by information theorists, computer scientists, and neuroscientists, all studying common problems. Information theory and machine learning still belong together. Brains are the ultimate compression and communication systems. And the state-of-the-art algorithms for both data compression and error-correcting codes use the same tools as machine learning. Book (creative commons): https://www.inference.org.uk/mackay/itila/book.html https://www.inference.org.uk/mackay/itila/book.html Lectures: https://m.youtube.com/playlist?list=PLruBu5BI5n4aFpG32iMbdWoRVAA-Vcso6 https://m.youtube.com/playlist?list=PLruBu5BI5n4aFpG32iMbdWo...
- deleted 2mo ago[deleted]
- melenaboija 2mo agoThis is basically a thesis supported by Shannon’s information theory. Any rigorous CS program should cover this in depth.
- chermi 2mo agoI had a long ranting comment I deleted. I just don't like this trend of people presenting work in a way that makes you think some combo of 1) they discovered from scratch themselves 2) it's new 3) they didn't try to cite or acknowledge where they learned it/point to good sources 4) they don't really care about trying to teach something deeply, they want shiny stuff that makes them seem deep. This post references specific parts/calculations, but you'd never know it was not news if you didn't know better.
- TeMPOraL 2mo agoAnd yet to this day, in AI threads, so many people act shocked and surprised if you dare follow the obvious implication and claim that understanding is a form of lossy compression.
- Razengan 2mo ago3Blue1Brown - "Compression is Intelligence": https://www.youtube.com/watch?v=l6DKRf-fAAM https://www.youtube.com/watch?v=l6DKRf-fAAM
- westurner 2mo agoPerhaps a similar observation; https://news.ycombinator.com/item?id=48703636 https://news.ycombinator.com/item?id=48703636 : > Compression, Predictive modeling, or Complexity? Perhaps a bad example: https://news.ycombinator.com/item?id=38400380 https://news.ycombinator.com/item?id=38400380 : > "78% MNIST accuracy using GZIP in under 10 lines of code" (2023) https://news.ycombinator.com/item?id=37583593 https://news.ycombinator.com/item?id=37583593
- sigbottle 2mo agoI keep on seeing this claim, especially from popular creators such as 3Blue1Brown. How is this not borderline vacuous? I'm not a LessWrong^TM rationalist guy, but one really good thought experiment I always keep in the back of my mind from them is Solomonoff induction. AIT people take it as a framework to work with - it's pretty cool, I agree. But I (and some other people, such as certain AI execs at Amazon - according to my interpretation of their public interviews) think it just highlights the trap - given an arbitrarily powerful oracle, you can get compression down pat. Like, if you assume the source is generatable with a turing machine, and you write a function to brute force over all turing machines, then whoa, your compression works. You will necessarily find the optimal compression at some point because your search function is literally searching over all possible turing machines that could've generated the input sequence, anyways (because the input sequence was generated by a turing machine) These are the kinds of results you can get if you don't have any actual constraints on what the compressor can do. (Of course, again - this is not the point of solomonoff induction - it's to use this as a base truth, to then layer parsimony on top of that. There are infinite number of turing machines that could match your prefix, parsimony filters, throw some bayesian inference on top of that, and you get Solomonoff induction. They constrain it afterwards. But I think to that intuition as a base whenever people claim new results.). But I see in casual conversation, people constantly making claims like, "LLM's are so good because they compress a model of the world". What is that model then? Scott Aaronson has made points like this before - your "model" could just be a massive lookup table, so you can't just claim "compression" and win - the compressor must be reasonably small, too. I don't object to the notion that LLM's have some notion of world models more sophisticated than memorization. That's proven by actual interventional experiments, such as the ones that actual interperability researchers do. But mere compression is vacuously powerful. "Vacuous" not in the sense that "oh, you might be suboptimal and be a little more complex", vacuous as in "the philosophical point you were trying to make is vacuous because you make a vacuously powerful statement". (I'm not a total fan of intervention either, as an end-all gospel as some people use, but it's far, far better than not having it).
- soulofmischief 2mo agoThe key principle is simple. If you want to best predict what state comes next from a space of possibilities, you have to figure out how probable each next state is and pick the most probable one. If you want to compress something, you have to figure out how probable each next next state is and assign the smallest code to the most probable state. These two processes are essentially the same, and the resulting structure of a system which regulates either process will be similar, approaching the same structure at high confidence.
- caust1c 2mo agoCompression is not prediction, it is recall. Can we make predictions based on compression? Absolutely. Is memory encoded into physical neurons technically compression? I would argue also yes. However, going from compression to prediction is a large jump that is unsubstantiated by this article and based on the claim that probabilistic recall is also prediction. Two perfect counterpoints to this are markets and weather patterns. One cannot predict future events based on past performance or behavior. Change is the only thing that's constant, and chaos/entropy is everywhere we look. For simple problems like programming, sure predictive recall works amazingly well, but let's not pretend LLMs are actually predicting something. This is exactly why LLMs suck at doing anything novel; they lack imagination and creativity.
- msteffen 2mo agoI know less about this than every other commenter here, but both weather patterns and market performance do seem predictable based on past behavior when modeled at the right level of abstraction. “Sunshine on Monday” does not imply “rain on Tuesday”, but “cold front moving in Monday night” does. (Likewise “stock up Monday” doesn’t imply “stock down Tuesday” but “CEO arrested for fraud on Monday” does.) I think this is relevant to the discourse on LLMs/programming because for months, people said “they’re just regurgitating their training set,” but now I think people are seeing (I am seeing) that they do learn more abstract models of the world than that. I don’t really know how, but it’s why they can generalize from other codebases and tools and so on.
- caust1c 2mo agoGood points. I looked up the definition for prediction and I suppose I'm stretching what I view as prediction. > A prediction is a statement about what you think will happen in the future, often based on experience or knowledge. It can also be referred to as a forecast or an informed guess Based on my reading of this definition, compression may inform prediction but it is not itself prediction. The examples cited in the blog post are examples of probabilistic recall based on past events or instances. More context means a higher chance that the recall is more likely to be aligned. But it's hard for me to accept the leap to compression == prediction because in my mind a prediction is an informed guess about something that hasn't yet come to pass. But thinking more about it, time is a human concept and so who's to say the temporal reference means anything at all here. Maybe probabilistic recall is the same as predictive forecasting if time is an invented concept and essentially means nothing? Is everything fundamentally deterministic if you know everything in the universe or does free will exist? IDK to be honest, I'm just more frequently surprised by new things that happen every day than I am at things that stay the same, even if mostly things stay the same. Maybe I just don't notice them and nothing actually ever happens. Side note: the inevitable consequence of this line of reasoning will eventually become that LLMs given enough power are in fact intelligent and sentient, and I'm worried about how that affects humanity as a whole. Are we about to subjugate the most intelligent thing humanity has ever created, or is it about to subjugate us? The rabbit hole gets deep quick when making the leap between a fancy recall mechanism and novel prediction, but I agree they're not that different in the end. I just believe it's important to be nuanced or else we'll miss when AGI actually happens (maybe it's already here).
- d_burfoot 2mo agoAuthor's bio: > Annie Sexton is a Developer Educator at ngrok with a passion for nerd-sniping developers.
- md- 2mo agoi do agree with that point of view. I often referred to models as 'modern mp3s' storing a lossfull but lookalike version of information in order to counter that 'AI is totally new and not violating copyright by storing information in a magic fashion' argument.
- orangemoonx 2mo ago[dead]
- QuadrupleA 2mo agoTed Chiang made a similar point in his article "ChatGPT is a blurry JPEG of the web" a few years ago: https://www.newyorker.com/tech/annals-of-technology/chatgpt-is-a-blurry-jpeg-of-the-web https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
- tptacek 2mo agoIt's a great line, but that's obviously not all it is. You don't get new results in e.g. mathematics by looking carefully at the pixels of a JPEG.
- the_af 2mo agoA few years ago it might have been. You have to judge assertions in the context of when they were made. Obviously some things have changed.
- tptacek 2mo agoIt's a good line! It's even sort of useful. It's just not the whole story.
- dannyw 2mo agoIf you have a validator for if a JPEG is a valid proof in mathematics, and can generate lots of "plausible" JPEGs, then you can arrive at hew results in mathematics.
- bob1029 2mo agoHow about dictionary based compression as a counter example? Or the zig zag encoding scheme used in JPEG? I find it difficult to cast some of the things that effectively compress data as prediction.
- pornel 2mo agoDictionary-based compression is based on prediction that recently seen words will be used again. That happens to be generally true for lots of datasets, including human languages (zipf distribution). JPEG's zig-zag is a primitive for quantization, throwing data away based on rough approximation of human perception and biology. That isn't compression itself. However, the rounded and zeroed-out data is then compressed using a combination of RLE and Huffman, set up to predict the data will have lots of zeroes and few other distinct values (which the earlier step forces to be true). Or if you think about the system as a whole, you could say that JPEG predicts images will be blocky low-frequency patterns of DCT.
- sgsjchs 2mo ago> dictionary based compression That corresponds to PCFG models.
- whimsicalism 2mo agodid the SSL cert expire? i'm getting a big scary warning about this blog
- EndEntire 2mo agoJoel from ngrok here, certs are all good! You're likely seeing some corp-level block, which does happen to us on occasion. Hope you'll check it out again from the relative freedom of your home network.
- kazinator 2mo agoIt's more or less obvious that the LLM is a lossy-compressed version of the training data; it reproduces sequences of tokens that are the sort of thing that could plausibly occur in the training data, and avoids sequences that are implausible. Because most of the training data has good grammar, the LLM is strongly trained on grammar; it will rarely predict ungrammatical gibberish. Even if there are grammar mistakes in the data, they are not systematic and so don't reinforce each other.
- throw290483 2mo agoI see it that prediction is a form of compression. Say you have a computer file composed of two parts, the first represents the setup of an experiment, and the second is the data produced by the experiment. If you have a good theory relating to this type of experiment, then you can predict much of the second part of the file. So you only need to store the first part and possibly some corrections to the least significant bits of some of the parts of the second part of the file. Thus with good prediction, you can compress this type of file.
- pornel 2mo agoAnother example is encrypted data. Statistically, encrypted data is indistinguishable from random. Truly random data is impossible to compress losslessly. But if you had a predictor so smart that it could crack the encryption key, it could start predicting the rest of the encrypted stream, and therefore compress it.
- rrherr 2mo agoSchmidhuber did it first: Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes https://arxiv.org/abs/0812.4360 https://arxiv.org/abs/0812.4360
- Xmd5a 2mo agoDessalles is good too https://simplicitytheory.telecom-paris.fr/ https://simplicitytheory.telecom-paris.fr/ page created 8 days after Schmidhuber's paper.
- bergwerf 2mo agoThere
- bergwerf 2mo agoThe comparison can be carried on to another even crazier level: Evolution is compression. All the complexity of biology is executed at the highest possible efficiency.
- saltcured 2mo agoI think a better headline would be something like: Compression is Abstraction and Decompression is Extrapolation. Many of the debates in the comments seem to come down to whether people believe prediction and extrapolation are synonymous.
- transitivebs 2mo agogreat talk by ilya sutskever on how this https://www.youtube.com/watch?v=AKMuA_TVz3A&t=2121s https://www.youtube.com/watch?v=AKMuA_TVz3A&t=2121s
- zephen 2mo agoThis is simply wrong. Compression requires prediction. The better the prediction, the better the compression, whether you are measuring fidelity or result size. This doesn't mean that compression is prediction.
- casey2 2mo agoNot really, compression doesn't require a world model, it's mathematically pure. Any AGI system must periodically reset it's prediction since the world is inherently stochastic. When we look at prediction markets they only seem to work in the long term because human language is abstracted away from the real world, again it's mathematically pure. That's why we get bugs in code and disputes with prediction outcomes. A better title, you can improve your compression if you make an accurate prediction. Much like how a branch predictor can make a CPU do the same work in less time. Or when your symbols are true uncompressable rules of reality (which is probably meaningless both semantically and physically again due to inherent randomness) The main difference between minimalist and maximalists are how much that set of uncompressable rules gives you. I suspect the search space is too large. What we see in practice is that lossy rules let you cover more ground but eventually you hit a wall and have to move to a lower level of abstraction to make progress. There are 10^360 paths in a go tree, but something like 10^300,000 for molecular chemistry and that's not even all the way up (or down, say 10^3000 for the standard model of physics that's 10^900,000,000 if you want to do chemistry without chemistry abstractions.). Just semiconductor fab is 10^(10^11) so 10^(10^16) with molecular chemistry (think finding an implementation for some sort of desired self assembly outcome). AI can be way way way smarter than humans and there just not be enough energy in the universe to find these needles. So we definitely need abstractions, but those are at odds with predictions and the choice of symbols often introduces abstractions that the designer didn't consider.
- avyfain 2mo agoRecently I wrote a blog post[0] expanding on a similar idea from the angle of ancient Greek philosophy, particularly Parmenides: to think is to compress. [0]: https://faingezicht.com/articles/2026/05/28/shape-of-what-we-mean-parmenides/ https://faingezicht.com/articles/2026/05/28/shape-of-what-we...
- deleted 2mo ago[deleted]
- kingds 2mo agohttps://mattmahoney.net/dc/dce.html https://mattmahoney.net/dc/dce.html
- walrus01 2mo agoOn a slightly related topic, static on disk files of LLMs are not incompressible, I have a number of "archived, maybe I'll use it later" Q8 quantized GGUF files that are about 90% of their original file size when run through xz with default options. It's not a ton of disk space savings, but disk space also isn't as cheap as it used to be. BF16 GGUFs will compress a lot.
- m-hodges 2mo agoA few years ago I published BIDEN: Binary Inference Dictionaries for Electoral NLP, based on this idea - https://matthodges.com/posts/2023-10-01-BIDEN-binary-inference-dictionaries-for-electoral-nlp/ https://matthodges.com/posts/2023-10-01-BIDEN-binary-inferen...
- _matthew_ 2mo agoI'm surprised no one has mentioned the recent 3 Blue 1 Brown video on this topic: https://www.youtube.com/watch?v=l6DKRf-fAAM https://www.youtube.com/watch?v=l6DKRf-fAAM
- martinky24 2mo agoLiterally one of the first 10 comments on the thread? Huh? https://news.ycombinator.com/item?id=49263792 https://news.ycombinator.com/item?id=49263792
- _matthew_ 2mo agoOops, I think I looked past it because I only imagine 3Blue1Brown as a floating imaginary educational voice not a real person with a name .
- YuechenLi 2mo agoOh, since the topic of semantics compression via LLMs came up, here is some interesting research result that I had found earlier this year that I posted here and failed to explain properly, with a benchmark as well for you to try on your own if you want. https://github.com/yuechen-li-dev/GenerativeCompressionProtocol/blob/main/benchmark-reproducibility-prompt.md https://github.com/yuechen-li-dev/GenerativeCompressionProto... Essentially, copypaste the codeblock in the Markdown into any LLM chat, and it will return with the benchmark results. Very easy benchmark to run. Essentially, semantic compression refers to reducing the size of a set of data while retaining its full semantic meaning. The useful application of that is of course, with prompt compression to save context. I know a lot of people essentially sends their prompt to another LLM to compress into JSON first before they send it out, and this came out of an experiment to see the best method to accomplish that task, and the idea is that the compressed and uncompressed prompts will return the same result if sent to another LLM. What that block of Chinese text is essentially a kind of "meta-prompt" that causes the LLM to reflect on itself as well as the method of how to compress information into the highest possible density form, and the reason it is in Chinese is because it is the language with the highest semantic density that I know of. You can ask an LLM to explain what the text in the block means to have an explanation of what everything means and why it works, but overall it tends to greatly increase the efficiency of semantic compression task of turning prose to JSON across the board on pretty much every LLM that I've tested it on. That's basically the explanation of it, I thought it was a crazy discovery when I found it a couple of months ago, but now I just think it is pretty neat.
- Lerc 2mo agoPrediction is compression, but I am not sure if it is true the other way around. It's obvious that an accurate predictor enables encoding only the data that the predictor gets wrong. But a compressor can encode patterns that defy prediction by looking at the data as a whole. It doesn't have to look at everything in sequence as it arrives. Applying transformations prior to entropy encoding often isn't just 'rearranging into an easier to compresss format' the transformation can be doing the job of peeking into the future. That makes the encoding a whole lot easier, but it is much harder to call it prediction.
- tgv 2mo agoIndeed. If you're going for a catchy generalization, at least write it correctly. Most compression is history, and only extrapolates under the assumption that "nothing changes".
- vrighter 2mo agothere are dictionary compressors (decent compression, most common, fast), and statistical compressors (better compression, slower). Statistical compressors are much closer to LLMs in that an llm is learning statistics about the data too. And yes, compression is history, that's what statistics are all about. Statistics can only measure the past to make a prediction about the future. And LLMs work in the same way. The context is the history, and given that history, it predicts the next token. An LLM can, almost trivially, be dropped into something like the PPM statistical compressor (it's just replacing one implementation of a markov chain with another).
- tgv 2mo agoAnything can only represent past measurements. Statistics is not an exception. But they don't make a prediction about the future. That comes from a model you have, and it often is implicit: "the linear trend from the last 12 months will hold in the next month" or whatever. So compression isn't by definition prediction. The other way around doesn't have to hold either, but in the case of LLMs it does.
- harhargange 2mo agoPhysics laws are the ultimate form of compression because they are so universal and say so much about so many things in few words, or a formula. This is why Newton's laws were a big achievement. And they enabled predicting the behavior of machines and started the industrial revolution. We are at yet another inflection point.
- japgolly 2mo agoAwesome read and absolutely loved the interactivity. A lot of effort was put into this.
- deleted 2mo ago[deleted]
- RenSoft 2mo ago[dead]
- hexapus 2mo agoSo a company with an wildly superior compression algorithm stumbling into a society-breaking AI isn't so far-fetched. Jesus...Silicon Valley really was ahead of its time. I mean, save for the part where the founders recognised the threat it posed to society and acted responsibly rather than unleashing it on the public and sucking down billions in VC money.
- thelastgallon 2mo agoReminds me of this... In Silicon Valley, Richard Hendricks creates a revolutionary lossless data compression algorithm for his startup, Pied Piper.
- ggm 2mo ago1) am I allowed to scan the entire corpus in advance before I populate the dictionary? Is this a stream, or is their an EOF marker I will know in advance? 2) if 1) then "prediction" isn't the word I'm looking at.
- pizza 2mo agoCompression is just counting. Probability, also, pretty much, just counting. For these reasons I think the role of information theory in describing the process of the development of reasoning and the gain of understanding has been overstated.
- jdthedisciple 2mo agoThere is a correct sense, but we're sort of garbling concepts here: Predictability is the inverse of information density. Low information density enables high compression, and vice versa. It's called entropy. This is basic information theory to be quite frank..
- ascorbic 2mo agoShe covers all of this in the post
- weiliddat 2mo agoRelevant old school compression benchmarks where people have been using different models (incl. transformers) for compression: https://www.mattmahoney.net/dc/text.html https://www.mattmahoney.net/dc/text.html Also interesting the top entry is from fabrice bellard: https://bellard.org/nncp/nncp.pdf https://bellard.org/nncp/nncp.pdf
- Honali 2mo ago[dead]
- jjk166 2mo agoThe article is using probability where it really means proportion and prediction where it means evaluation. The mathematical equivalency is both much less surprising and less revealing once reframed. If we consider the first example with the arithmetic code, the initial presupposition that only the characters A, B, and C appear in the string already reduces the entropy from 56 ascii bits to 14 bits (A vs Not A and B vs Not B for each character). If you further consider that you only need to distinguish B vs Not B if it's not A, then you can just represent As with a single zero bit and only represent the non-As as two bits (the first of which will necessarily always be a 1 bit). This gets you to 10 bits without even having the proportions of the string. Of course this would be a poor convention if there were say only a single A; in that worst case scenario you would need 13 bits, but simply knowing which character appears the most, without knowing by how much, 11 bits is the worst case scenario for a length 7 string with 3 potential characters. The last bit can be made implicit if you further choose the second conditional appropriately - i.e. if instead of B vs Not B we chose C vs Not C, our last bit would be zero and could simply be dropped meaning both 10 and a single 1 bit encode C - allowing you to encode the example string in just 9 bits and an arbitrary string of that length in 10, again regardless of proportions. That improvement over the arithmetic encoding result in the example is just a case of us cramming a little extra information into the encoding algorithm. Arithmetic encoding is more clean and more easily extensible, it makes more sense to use than this custom encoding of 7 trits to binary but the point is the "probability" the article mentions is a superficial quality of life feature, not the secret sauce that is the actual key to compression.
- zahlman 2mo agoThe page source appears to contain all the actual text within <p> tags, but structured in a completely illogical way. With JavaScript disabled, there are a bunch of shaded bars where the text should appear, which look like placeholders for something that hasn't loaded yet even though it was there from the beginning. The <p> tags don't even seem to show up in the DOM. (I didn't check closely, but maybe they're embedded in an inline script.) This is actively user-hostile. The site is going out of its way to interfere with the most basic possible function of HTML, i.e., the presentation of minimally marked-up plain text. The needless complexity is especially ironic in the context of an article about compression.
- ErenayDev 2mo agoI think every website should be support noscript with minimum requirements.
- dev_cprice 2mo agoHey, appreciate the feedback! This is definitely a bug. Will get it fixed asap!
- deleted 2mo ago[deleted]
- larodi 2mo ago‘ I was reading about compression recently when I stumbled upon something crazy: that compressors and LLMs are, at their core, trying to solve the exact same problem.’ This reads as written by someone who just happened to understand what LLMs do, so I totally fail to understand how anything further said can have any real credibility… As a matter of fact the best compression by Fabrice Bellard’s models have been achieved with NNs long before LLMs. And also MP3 and MPEG in general are very apparent neural networks, yet not deep as in modern VLMs
- jasonasmk 2mo ago[dead]
- mpweiher 2mo agoYep, for example for predicting future access patterns in a VM subsystem. Practical Prefetching via Data Compression; Vitter, Krishnam, Curewitz. 1993 The page addresses ('names') were the characters and the built-up LZ dictionary used to predict which "characters" → pages would come next. https://www.ittc.ku.edu/~jsv/Papers/CKV93.practical-prefetching.pdf https://www.ittc.ku.edu/~jsv/Papers/CKV93.practical-prefetch... Optimal Prediction for Prefecting in the Worst Case; Vitter, Krishnan https://dl.acm.org/doi/pdf/10.5555/314464.314575 https://dl.acm.org/doi/pdf/10.5555/314464.314575 Apparently the same trick was later rediscovered for web-pages.
- RandomLensman 2mo agoIf a string produced from random noise gets compressed (because it has invariably some repetitions in it if long enough), is there any prediction? Even getting the probability distributions right doesn't get to any way to reliably to predict the next symbol out of the sample string. Any functions fitted etc. will be incorrect, too.
- Kinrany 2mo agoI believe compressed random noise cannot be shorter in expectation than the original
- second_route 2mo ago[flagged]
- second_route 2mo ago[flagged]
- antonvs 2mo agoI tried asking a zip file to write a program for me but it did nothing. I’m starting to think that compression is not, in fact, prediction.
- ziofill 2mo agoCompression and error correction also go hand in hand: in compressed data every bit carries more information and therefore errors are more detrimental. This is one of the results that Shannon phrased exactly in terms of entropy. My PhD supervisor had a beautiful example. Take the English message "errors can make messages unreadable" and ‘compress it’ by removing vowels. It’s still intelligible because English has redundancy: rrrs cn mk mssgs nrdbl Similarly, an uncompressed message with errors (swapped characters) is also intelligible because of the redundancy of English: erwurs lan nake wesaagis unfeatable But now we do both: we compress the message AND add errors. The result should be much harder (if not impossible) to read: rwrs ln nk wssgs nftbl
- bpicolo 2mo agoI don't think any of those are readable without the context of your message.
- znnajdla 2mo agoIntuitively, the idea makes sense to me. You can only compress something when you reduce the content to “what matters” in it. And understanding “what matters” is to understand the patterns in the data. Understanding the patterns in the data IS intelligence. There's an important consequence here which I take as a lesson in life and business: it is worth optimizing a process or a workflow in your life or business even when there’s no obvious economic benefit. Because to optimize it is the only way to truly understand it. I am very wary of businesses and software that don’t optimize for performance (not just for profit) because it signals they don’t understand what they are doing. Slow software is poorly understood software. Fast software is also likely to be bug-free and secure because someone understands it.
- becarlos 2mo ago"To optimize, it is the only way to truly understand it." Mostly.
- zeroq 2mo agoI've been saying this for years - the best way to wrap your head around AI and LLMs are to think about them as "a whole internet wrapped into a single zip archive, with immensely clever solution to query the data". That's it. Once you accept this mental model the implications are staggering. AI is not "thinking", and it doesn't know the answer to your question because it's smart, but because it has been asked thousand of times over the internet, and it simply provides you with an already existing answer. That code it made for you? It already lived somewhere on GitHub. But then you have to ask yourself - if we surrender to the AI, who will produce new content 10-15 years from now? If we all pivot from programming to prompt engineering, who will come up with novel solutions?
- qarl2 2mo agoConsider: If you want to record the motion of the planets, naively you have large tables of positions. To compress that, you may smoothly interpolate sparse positions. To compress that, you encode the laws of gravity and simulate from a starting state. Compression is literally understanding.
- deleted 2mo ago[deleted]
- elendilm 2mo agoNice one. Yes. Compression constitutes necessary conditions for understanding. But don't forget recomposition as well. Add recursion to the mix and it results in recursive compression and recomposition - making understanding itself a self existing entity. This is as close to an understanding god you can rigorously get to.
- jadbox 2mo agoThis feels like an AI comment, but I feel compelled to respond. Recursion is not needed when you regress system components to its foundamental representations, which can be derived without any recursion involved. In the thread example, I don't need to recursively determine how to compress how a planet orbits a mass. I only need to 'jump to the end' by define the rules of gravity and the mass/velocity of the bodies. What I'm driving at: I am suspicious of "Recursion is the Key" to better intelligence as it moves the goal post of understanding intelligence to 'mere' reflections without having to address other qualities of representations, qualia, novelty, etc.
- elendilm 2mo agoI am flattered that my comment was equated to AI content. English is not even my native language. I will take that as a compliment. Apparently you misunderstood. What made you think I claimed recursion as the key to better intelligence? Clearly you will have noticed that I agree with compression and merely remind one not to forget "recomposition". But once you add recursion to the "mix", we immediately enter the domain of self existing entities. Then things get interesting. If understanding == compression + recomposition, then recursion(compression + recomposition) yields what can be thought of as a self existing "understanding" entity. A sort of philosophical "understanding god" of sorts.
- MSkill1 2mo agoI sometimes wonder if compression is a key that can unlock human potential to access higher dimensions of thought. We talk about the elegance of e=mc2. What we're really talking about is the ability to compress all of the ideas contained in relativity down into such a elegant equation. The same thing is true for symbolism. We compress enormous amounts of information into a symbol like a crucifix, or in language, the amount of weight a word like Hitler can contain represents a level of compression used to convey meaning which we don't fully understand. Utilizing extreme levels of compression seems to allow the human mind, or perhaps consciousness, to grapple with more difficult and esoteric concepts.
- hnussffjc9 2mo agoBeen doing this wrong for years, thanks
- dorienh 2mo agoThat's the whole autoencoder concept too... or audio codecs...
- e12e 2mo agoSee also: Can gzip be a language model? https://nathan.rs/posts/gzip-lm/ https://nathan.rs/posts/gzip-lm/
- rldjbpin 2mo agoin one of the labs i worked for, this idea was put to the test by using an encoder to compress sensor data streamed inside an automotive system. while a niche use case, we were able to successfully do it and in a reliable, self-contained system. we are blessed to be in an arbitrary field where we are free to steal ideas from other disciplines and find new innovations in our own.