5 ms·
This is the thesis behind the "Information Theory, Inference, and Learning Algorithms" course that was taught at Cambridge University. > Why unify information
by farfatched 2mo ago
This is the thesis behind the "Information Theory, Inference, and Learning Algorithms" course that was taught at Cambridge University.
> Why unify information theory and machine learning? Because they are
two sides of the same coin. In the 1960s, a single field, cybernetics, was
populated by information theorists, computer scientists, and neuroscientists,
all studying common problems. Information theory and machine learning still
belong together. Brains are the ultimate compression and communication
systems. And the state-of-the-art algorithms for both data compression and
error-correcting codes use the same tools as machine learning.
Book (creative commons): https://www.inference.org.uk/mackay/itila/book.html https://www.inference.org.uk/mackay/itila/book.html
Lectures: https://m.youtube.com/playlist?list=PLruBu5BI5n4aFpG32iMbdWoRVAA-Vcso6 https://m.youtube.com/playlist?list=PLruBu5BI5n4aFpG32iMbdWo...
- deleted 2mo ago[deleted]
- melenaboija 2mo agoThis is basically a thesis supported by Shannon’s information theory. Any rigorous CS program should cover this in depth.
- chermi 2mo agoI had a long ranting comment I deleted. I just don't like this trend of people presenting work in a way that makes you think some combo of 1) they discovered from scratch themselves 2) it's new 3) they didn't try to cite or acknowledge where they learned it/point to good sources 4) they don't really care about trying to teach something deeply, they want shiny stuff that makes them seem deep. This post references specific parts/calculations, but you'd never know it was not news if you didn't know better.
- TeMPOraL 2mo agoAnd yet to this day, in AI threads, so many people act shocked and surprised if you dare follow the obvious implication and claim that understanding is a form of lossy compression.
- kazinator 2mo agoShorter description isn't understanding, let alone of it is lossy. When you shorten a description in a lossy way, you are deciding a priori that some differences in the object don't matter, and it's not because you understand the object, but because it serves your goal of shortening the description.
- Dylan16807 2mo agoOr you actually do understand it. You can't just assume smaller is better but it often is. And very often it's more information-dense.
- kazinator 2mo agoYou can compress syntax, losslessly even, with zero understanding of its semantics. Zero understanding not only imbued into the compressor/decompressor, but even the designer of the compressor doesn't require understanding the semantics. Actually, even of the syntax. A compression program can compress a book written in a language that the author of the program doesn't understand, on a topic he knows little about.
- Dylan16807 2mo agoFinding common characters and building a list of words is a low level type of understanding. Doing it better does actually start directly representing syntax patterns and that's a less-low level of understanding. I think "losslessly even" is the wrong way to think about it. Lossless compression often requires less understanding than high quality lossy compression. If you can do a lossy compression that correctly decides what details are unimportant, that's a good sign of understanding.
- skinfaxi 2mo ago> I think "losslessly even" is the wrong way to think about it. Lossless compression often requires less understanding than high quality lossy compression. If you can do a lossy compression that correctly decides what details are unimportant, that's a good sign of understanding. This is the crux and reminds me of things like mp3 that exploit the nature of human hearing being limited to a frequency range.
- exe34 2mo agoReminds me of Stephen Wolfram "discovering" things in the sense that other people would say "today I learnt".
- jagged-chisel 2mo agoI find it bothersome that language works this way. You can spend your whole life discovering things that are well known by the rest of the world. But the minute that you mention to a large group that you “discovered” it, suddenly you’re taking credit for discovering it for all of mankind.
- deleted 2mo ago[deleted]
- chermi 2mo agoWolfram was my example in my original rant!
- mpalmer 2mo agoI'm glad to see someone feels similarly. There is nothing wrong with ignorance, but there's no excuse mistaking learning for invention. Especially from someone bearing the title "Developer Educator" I don't think it's the case here, but worth noting too that LLM-written blog posts adopt this tone seemingly by default. Never the least bit of surprise, wonder, doubt, or frustration to get in the way of the steady staccato beat of metaphors, conclusions... and three-item lists.
- bch 2mo ago> I'm glad to see someone feels similarly. There is nothing wrong with ignorance, but there's no excuse mistaking learning for invention. Especially from someone bearing the title "Developer Educator" >> a Developer Educator at ngrok with a passion for nerd-sniping developers. Maybe more the latter than former...
- farfatched 2mo agoI don't think this is a fair critique. The author of the post uses standard terminology like entropy coding and arithmetic coding, and cited a paper "in 2023, Google DeepMind released a paper arguing that language modeling and compression are two views of the same thing" which discusses it further. This blog post is great. Well explained, and clearly took a lot of effort. I don't interpret it as them claiming to have to discovered it independently.
- bonoboTP 2mo agoCiting 2023 makes it seem like this is newer than it is. Compression, prediction and intelligence have long been known to be deeply connected.
- beagle3 2mo agoYou expect every blog post to find the earliest relevant paper to cite, just so one could look at the year (without reading said paper - which would have made clear that the connection isn’t recent) to assess novelty? I don’t think that’s reasonable. It’s a blog post. If it was, say, a peer reviewed paper by Hinton or LeCunn that fails to cite Schmidhuber, that would be reasonable criticism in my opinion. (Spoiler: they fail to cite him)
- hnfong 2mo agoWhy would blog posts not be subject to such criticism? Either the author knew of prior work that argues the same thing and they ignored it, or they didn't know. And if one writes a 1000+ word article premised on this idea, wouldn't one be presumed to know at least in which century the idea originated from? Arguably these kind of blog posts should be more subject to such criticisms, because the blog posts purport to "teach" the general public about a concept in an authoritative tone (or at least the author seems to pose as knowledgeable in the subject), while for academic papers, everyone who actually reads the paper knows where the ideas came from anyway and it's mainly an issue of attribution (and maybe about fairly distributing the citation count...)
- bonoboTP 2mo agoI'm of two minds here. The pro is that the "you could have invented this" walkthrough from first principles is more engaging than "and then so and so introduced this term in 1972 and the definition is such and such". This style is a reaction to that boring and dry teaching style and tries to push towards what eg Feynman pointed at in the Brazil critique. The con is that you don't get to understand and see any of the history of the ideas or even the ballpark when it was discovered, you attribute it to the blog mentally and you don't know what is how new or old and can't reference it properly when talking to others.
- detourdog 2mo agoDisconnecting idea development from it's historic development is a disservice to the audience that may want to dig deeper.
- wpietri 2mo agoIf only there were some sort of way for a reader to dig deeper on a topic without a writer having to spoon-feed them the entire history of everything! It's wild to me what people here expect out of something they got for free and that was offered as a gift.
- detourdog 2mo agoMentioning any idea disconnected from its roots may not provide the terms needed to search. Snarky replies always appreciated
- wpietri 2mo agoYour theory is that anybody who writes anything is obligated to make sure you can find any related information with one Google search? Again, to me that looks like wanting to be spoon fed. I have no idea why you think the world owes you endless 101-level discourse, but I hope you recognize you're setting yourself up for equally endless disappointment. If you take a little responsibility for your own education, you'll be happier.
- hellohello2 2mo agoYou're reading this the wrong way I think, citations aren't given because its obviously a pedagogical article about well established stuff. Much like you wouldn't give citations in a blog post explaining calculus.
- detourdog 2mo agoOne could give citations regarding calculus it's pretty interesting. Since it was done twice by both Newton and Liebniz. There must have been cultural developments in the 1660's that demanded calculus be invented.
- hellohello2 2mo agoYeah I agree, I was just saying I really doubt the author of that blog post was trying to take credit for the ideas they present.
- detourdog 2mo agoYes there is chasm between not citing references and stealing ideas. I assume no malice and just want contextual explanations.
- teekert 2mo agoThe post says this is all part of gzip and LLMs, what are you saying? I’ve been using gzip my entire life. I read between the lines “this is common knowledge” throughout the piece. Throwing in some names and dates only makes this super clear story harder to read (and more like studying then the playful exploration this post was intended as).
- bonoboTP 2mo agoA short paragraph at the end on the origin of these ideas can be an easy way to dispel the misconception of potential beginner readers that the insights are novel.
- lumost 2mo agoI think you'll find that 95% of all academic presentations are telling stories out of other peoples work.
- chermi 2mo agoYes, and it's made pretty obvious with all the background section and references.. which btw I'm not advocating for such baggage in a blog post. Just something.
- jdthedisciple 2mo agoI'm glad you're pointing this out because not only are these old insights, but I'm also pretty sure I've seen variations of this blog post years ago on even HN already. The author acting as if they discovered this independently had me feel the exact same way. Kinda irritating and almost ... disrespectful? Not sure of the right words to describe it tbh
- vasco 2mo agoFabrice Bellard published this in 2023: https://bellard.org/ts_zip/ https://bellard.org/ts_zip/ > The ts_zip utility can compress (and hopefully decompress) text files using a Large Language Model. The compression ratio is much higher than with other compression tools. It's not only an old idea it's been totally done already.
- userbinator 2mo agothey discovered from scratch themselves If you followed the data compression scene in the 80s and early 90s, there were plenty of reinventions of LZ-ish and Huffman-ish algorithms (I also coded my own variant...), and people even tried to patent some of them, so at least for the basics I think it is something that many can discover independently; of course in these times, it's more likely they didn't.
- ascorbic 2mo agoThe first sentence says she came across it when reading about compression. I didn't read that as her claiming to have discovered the idea or that it was a new idea, I read it as "today I learned". I think somebody who was unfamiliar with how compression algorithms or language models work would find this an approachable and interesting introduction. Not everybody studied information theory.
- foo42 2mo agoPerhaps tangential to your point, but I often write blog posts (although finish and publish far fewer than I start) where I write about something as it has occurred to me, informed by things I've absorbed no doubt, but without specific research. In such cases I explicitly avoid searching out prior work as a) seeing that something is well discussed and explored can take away the motivation to explore (in the same way reading puzzle solutions before starting might), and b) to avoid having green shoots of ideas shaped by the current of existing consensus. Now that doesn't mean I don't come back after doing my own thinking to see what the more well developed literature of people cleverer than me, who've thought far longer than me think; I just don't want to snuff out my own exploration at the start. As I say most of these I never publish as I'm mainly using writing as a vehicle for thought, but when I do I'm never sure how to flag them. I don't want (imaginary, lets be honest) readers thinking I'm deluded into thinking I've found something new. I want to come up with a tag I can put on them which adds a pithy disclaimer card at the top or something so I feel more comfortable publishing them.
- jwr 2mo agoDoes everything have to be "news"?
- smath 2mo agoAh Sir David MacKay. I so respect him. Great explainer and speaker. He had built this text entry tool called Dasher [0] - that I'd heard him introduce at Princeton around 2006. It was basically an early language model that predicted which characters are more likely than others, given what you've already types and it would adjust the sizes of the available next characters based on their probabilities. [0] https://dasher.at/about/ https://dasher.at/about/
- farfatched 2mo agoHe really was fantastic, and prolific in multiple fields. He wrote https://www.withouthotair.org/ https://www.withouthotair.org/ (creative commons) and was the Chief Scientific Advisor to the UK Department of Energy and Climate Change. Dedicated to "to those who will not have the benefit of two billion years' accumulated energy reserves".
- jgraham 2mo agoI also went to a couple of his (fantastic) undergraduate courses, and have a huge amount of respect for him. That said, I think it's worth mentioning that Climate Change Without the Hot Air has aged pretty badly, and I'd be reluctant to recommend it to people who don't already have the background to understand what's aged well and what hasn't. The high level approach of making high level numerical estimates makes sense, but it dismisses solar energy in about a page due to assumed high costs. It turns out that even if you're David Mackay you can still be caught out by exponentials :) I notice now that the version you link has some inline updates pointing out how off the assumptions in this section were, but it seems to me that's not enough; you probably need to redo the entire analysis based on what we know today rather than trying to make purely local adjustments. On the other hand the point at biofuels are even more inefficient, and therefore a dead end even before you consider broader environmental impacts, are well made and something that is sadly not yet widely reflected in policy.
- 27183 2mo ago> biofuels are even more inefficient, and therefore a dead end Only if energy density doesn't matter. But it really does, though. Battery powered electric trucking? Dead end. Battery powered aviation? Dead end. Battery powered shipping? Dead end. [edit] Maybe there's some sustainable way to convert solar energy into sufficiently energy dense fuels that isn't biological, but so far it seems like seed oils or algae are probably the least bad?
- blahblahson 2mo agoBetter prediction being better compression is Shannon 1948, and the link to machine learning is MacKay 2003 at Cambridge.
- augment_me 2mo agoAs much as I want to, I sadly don't think Information Theory makes sense in this setting, and I really wanted to believe this. When Shannon made his theory of information, he was always dealing with informational representations on the abstraction level of bits. At Bell Labs, a lot of the work was on the compression of data for transfer over telephone wires. Entropy coding, later codexes like algorithmic coding, and all compression on this level assumes that you have a bit-based X, and you compress it. However, in deep neural networks, you are dealing with compression on different levels of abstraction. How do you decide what shared features a peacock and a palm tree have? At what scale should they be represented? How do you deal with invariance under affine transforms? Do you want to open the box of invariance under non-affine transforms? When you start looking at what it would mean to compress feature representations, you immediately get to the question of data. You realize that Shannon simply was given a form of a very low abstraction data and that information theory came out to handle data at this level, but it's not suited for the data representations of many higher level modalities. If you read Society of Mind by Marvin Minsky, which has aged well to about 80%, you can get the hint of the kind of abstractions that humans make and what would be needed to represent them, this is not representable in bits, you need to go to higher level shared features, and then you open all of the questions above as well as credit assignment, mutual information approximation, Fischer information between bayesians, etc.
- chacham15 2mo ago> How do you decide what shared features a peacock and a palm tree have? At what scale should they be represented? How do you deal with invariance under affine transforms? Do you want to open the box of invariance under non-affine transforms? The whole point is that the representation is learned. When you talk about various levels of abstraction, you're missing that all of these levels are representable with words and the relationships between them. That is verbatim what LLMs are optimized for. Interestingly, when you take an embedding, you do see that some transformations in embedding space actually hold which is quite interesting (e.g. tree + many ~ forest)
- augment_me 2mo ago1) I am talking about representations beyond language models and language embeddings. If you take for example image, video, audio, 3D-spatial DICOM or combinations like VLMs. If you ask a language model to make an image of a Begonia ferox leaf without training it with images as well, it will not be able to represent this. 2) Language is already a higher-order lossy compressed abstraction made by humans to communicate fast and fill out the left out information with a learned prior. If you train a model on language only, it will not have the opportunity to have a non-compressed representation to make its own abstraction from. 3) If you are LLM-pilled and believe that we will be able to reach arbitrary levels of precise informational representation using language only, and that all abstractions that we may ever want can live on every single embedding layer in an LLM, your argument is fair.
- ChuckMcM 2mo agoIt also helps explain to people that LLMs are as likely as bzip to develop "consciousness".
- sdenton4 2mo agoLet alone a sack of wet, self replicating protein! Just endless copying... How could it ever do anything more?
- ChuckMcM 2mo agoI suspect you're being snarky :-) but this is a really interesting question, and one that has had a lot of research done. I'm not current (I stopped following folks doing this research closely around 2019) but what we 'didn't' know about how brains work was still huge. Signaling levels, enzymes, the connectome, quantum effects, it is a really deep question. That said, once we do get a working idea of how it works, and can perhaps synthesize a brain artificially with proteins, it will inform us on the next steps for silicon realization of that.
- enneff 2mo agoI think the point is that tremendous complexity can arise from relatively simple mechanisms. That is what life is, at many levels. I’m not at all convinced that the current LLM approach will yield something we can broadly call consciousness but saying that it’s a simple concept and therefore won’t support consciousness is a specious argument imo.
- ChuckMcM 2mo agoI completely agree, tremendous complexity can arise from simple mechanisms. Gleick's Chaos is a really good introduction to that. I was talking about the article though, and the mechanisms currently used for training and inference in LLMs. Those mechanisms are mathematically precise (unlike Chaotic attractors) and as the author points out, achieve the same function as compressors do in a strict bit pattern minimization role. Sometimes tensor math is pretty complex, like the FFT and DCTs on JPEG compression, but with the same inputs you get the same results. And while a JPEG will never decompress to a different image than the one that was compressed in the first place, LLMs do not 'infer' token streams that haven't been trained in their training process. The big difference here is that if you imagine a JPEG compressor that compresses 100 different images into one 'chunk', you can see how to provoke it to produce any one of the images it previously compressed. And with a bit of creativity you can have it express different images in different parts of the resulting composite. FWIW I looked at patenting something like this for digital cameras to give them more "shots" space for a given amount of SD storage.[1] Given the way that models work in 'inference' mode (vs 'training' mode) you can't forward bias the result into the correct result when there are multiple forward results that have identical weights. It's the root cause of hallucinations, and you've lost information in the training phase that you can't then use to discriminate between the 'right' answer and an equally valid 'wrong' answer. [1] FWIW I could never recover enough state to insure that the image it regenerated was all of the same image you took. So you might get the street but one of the houses might be a house that was in a different picture you took. That kind of bug. Mostly arising out of the same kind of problem you have with using hashes to find documents, when you get a hash collision two documents have the same hash, so you don't know which one to return.
- gofreddygo 2mo agoI recently found Ross Ashby's papers on cybernetics and that led me finding out that his family published his entire indexed handwritten notes archive online [1]. Genuinely mindblowing. [1]: https://ashby.info https://ashby.info