5 ms·
This is covered in "Information Theory, Inference, and Learning Algorithms" by David MacKay ( https://www.inference.org.uk/itprnn/book.pdf https://www.inference
by rodlette 3y ago
This is covered in "Information Theory, Inference, and Learning Algorithms" by David MacKay ( https://www.inference.org.uk/itprnn/book.pdf https://www.inference.org.uk/itprnn/book.pdf ):
> Why unify information theory and machine learning? Because they are
two sides of the same coin. In the 1960s, a single field, cybernetics, was
populated by information theorists, computer scientists, and neuroscientists,
all studying common problems. Information theory and machine learning still
belong together. Brains are the ultimate compression and communication
systems. And the state-of-the-art algorithms for both data compression and
error-correcting codes use the same tools as machine learning.
* In compression, gzip is predicting the next character. The model's prior is "contiguous characters will likely recur". This prior holds well for English text, but not for h264 data.
* In ML, learning a model is compressing the training data into a model + parameters.
It's not a damning indictment that current AI is just compression. What's damning is our belief that compression is a simpler/weaker problem.
- retrac 3y agoCompression is abstraction.
- ultrarunner 3y agoI haven't read MacKay's work, so maybe this is naive, but I think the belief is that compression is a deterministic endeavor, whereas AGI may not be. The ability to recall information is incredibly useful; the ability to do so in convenient and flexible ways doubly so. The level of complexity being proportional to the GPT's ability to interpret means that compression is not necessarily a simple problem, like you say. However, if intelligence involves nondeterministic traits, i.e. something in the vein of creation of the present, AGI could be a significantly different problem to solve than compression. I think there's at least an intuition that this is the case, which explains the belief that compression is a simpler/weaker problem. As an aside, I'm currently unsure of my position on this.
- earthboundkid 3y agoLLMs are deterministic in principle. It’s like JPEG where it’s a lossy compression plus some deliberate injection of noise to add variety.
- a_cardboard_box 3y agoI'm convinced that deterministic AI will never fully replicate human intelligence. I will attempt to explain my reasoning. I can say with absolute certainty that I have subjective experience. If my behaviour is deterministic, I can imagine a philosophical zombie version of me that behaves the same but doesn't have subjective experience. That zombie, using the exact same reasoning as me, down to individual particle interactions, will (incorrectly) determine with absolute certainty that it has subjective experience and is not a philosophical zombie. Its reasoning, which is my reasoning, is thus flawed. Therefore, in order to believe my behaviour is deterministic, I must doubt my ability to reason. (Some claim that philosophical zombies aren't possible to imagine. But I can imagine them, so if that's impossible then my reasoning is still flawed just for a different reason.) I believe the same argument applies even if I assume my behaviour is non-deterministic but follows a probability distribution which is a function of the past -- the philosophical zombie has the same probability of reasoning that it has subjective experience as I do, and the rest of the argument is similar. Thus, I believe my behaviour is not only non-deterministic, but mathematically ineffable, and thus outside the realm of what computers can do. If you have subjective experience which works similar to mine, maybe you can follow my reasoning. If you don't, the argument may appear to be nonsense due to the ineffability of subjective experience.
- markisus 3y agoI don’t quite understand your zombie argument but imagine it would be quite easy to inject randomness into an AI agent.
- lossolo 3y agoWe are already doing it, as all LLMs are fully deterministic. So we are adding random seed to prompts and control temperature.
- 3cats-in-a-coat 3y agoActually, due to parallel execution they are not running deterministically. Even if the temperature is zero some randomness is left in. You can run a model deterministically, but this is incredibly slow.
- yreg 3y agoI’m unsure human intelligence is nondeterministic.
- sam_lowry_ 3y ago> that compression is a deterministic > endeavor, whereas AGI may not be. Fabrice Bellard famously made a compression tool using a Transformer [1]. So if transformers can be deterministic... why not AGI? [1] https://bellard.org/nncp/ https://bellard.org/nncp/
- earthboundkid 3y agoYes, this is a well known thing to anyone paying attention to information theory.
- DonHopkins 3y agoI love David MacKay's brilliant work on the Dasher text input system, which draws deeply from his work on information theory -- imagine Dasher integrated with an IDE and code search and Copilot and language model! "Writing is navigating in the library of all possible books." -David MacKay We just allocate more shelf space to the more probable letters. Why isn't Dasher built into every operating system and mobile phone? https://en.wikipedia.org/wiki/Dasher_(software) https://en.wikipedia.org/wiki/Dasher_(software) https://dasher.acecentre.net/about/ https://dasher.acecentre.net/about/ https://news.ycombinator.com/item?id=17105728 https://news.ycombinator.com/item?id=17105728 DonHopkins on May 18, 2018 | parent | context | favorite | on: Pie Menus: A 30-Year Retrospective: Take a Look an... Dasher is fantastic, because it's based on rock solid information theory, designed by the late David MacKay. Here is the seminal Google Tech Talk about it: https://www.youtube.com/watch?v=wpOxbesRNBc https://www.youtube.com/watch?v=wpOxbesRNBc Here is a demo of using Dasher by an engineer at Google, Ada Majorek, who has ALS and uses Dasher and a Headmouse to program: https://www.youtube.com/watch?v=LvHQ83pMLQQ https://www.youtube.com/watch?v=LvHQ83pMLQQ Another one of her demonstrating Dasher: Ada Majorek Introduction - CSUN Dasher https://www.youtube.com/watch?v=SvsSrClBwPM https://www.youtube.com/watch?v=SvsSrClBwPM Here’s a more recent presentation about it, that tells all about the latest open source release of Dasher 5: Dasher - CSUN 2016 - Ada Majorek and Raquel Romano https://www.youtube.com/watch?v=qFlkM_e-sDg https://www.youtube.com/watch?v=qFlkM_e-sDg Here's the github repo: Dasher Version 4.11 https://github.com/GNOME/dasher https://github.com/GNOME/dasher >Dasher is a zooming predictive text entry system, designed for situations where keyboard input is impractical (for instance, accessibility or PDAs). It is usable with highly limited amounts of physical input while still allowing high rates of text entry. Ada referred me to this mind bending prototype: D@sher Prototype - An adaptive, hierarchical radial menu. https://www.youtube.com/watch?v=5oSfEM8XpH4 https://www.youtube.com/watch?v=5oSfEM8XpH4 >( http://www.inference.org.uk/dasher http://www.inference.org.uk/dasher ) - a really neat way to "dive" through a menu hierarchy/, or through recursively nested options (to build words, letter by letter, swiftly). D@sher takes Dasher, and gives it a twist, making slightly better use of screen revenue. >It also "learns" your typical useage, making more frequently selected options larger than sibling options. This makes it faster to use, each time you use it. >More information here: http://beznesstime.blogspot.com http://beznesstime.blogspot.com and here: https://forums.tigsource.com/index.php?topic=960 https://forums.tigsource.com/index.php?topic=960 Dasher is even a viable way to input text in VR, just by pointing your head, without a special input device! Text Input with Oculus Rift: https://www.youtube.com/watch?v=FFQgluUwV2U https://www.youtube.com/watch?v=FFQgluUwV2U >As part of VR development environment I'm currently writing ( https://github.com/xanxys/construct https://github.com/xanxys/construct ), I've implemented dasher ( http://www.inference.org.uk/dasher http://www.inference.org.uk/dasher ) to input text. One important property of Dasher is that you can pre-train it on a corpus of typical text, and dynamically train it while you use it. It learns the patterns of letters and words you use often, and those become bigger and bigger targets that string together so you can select them even more quickly! Ada Majorek has it configured to toggle between English and her native language so she can switch between writing email to her family abroad and co-workers at google. Now think of what you could do with a version of dasher integrated with a programmer's IDE, that knew the syntax of the programming language you're using, as well as the names of all the variables and functions in scope, plus how often they're used! I have a long term pie in the sky “grand plan” about developing a JavaScript based programmable accessibility system I call “aQuery”, like “jQuery” for accessibility. It would be a great way to deeply integrate Dasher with different input devices and applications across platforms, and make them accessible to people with limited motion, as well as users of VR and AR and mobile devices. https://web.archive.org/web/20180826132551/http://donhopkins.com/mediawiki/index.php/AQuery https://web.archive.org/web/20180826132551/http://donhopkins... Here’s some discussion on hacker news, to which I contributed some comments about Dasher: A History of Palm, Part 1: Before the PalmPilot (lowendmac.com) https://news.ycombinator.com/item?id=12306377 https://news.ycombinator.com/item?id=12306377
- WhitneyLand 3y agoI don’t think the statement is annoying to some because they see it as a damning indictment of AI. It’s because it’s not literally true. If it were literally true, a zip file would be ChatGPT, but it is not. The paper was named in an intentionally provocative way. What they really mean is, at the heart of these transformer models, compression is the fundamental principle or goal. This trend in naming papers has evolved over time. It didn’t always used to be this way. Some would say a provocative title is thought-provoking and drives a wider interest. Others would say it’s a research paper and things in it should be literally true. edit to add direct link to paper: White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? https://arxiv.org/abs/2311.13110 https://arxiv.org/abs/2311.13110
- seanhunter 3y agoCompletely agree. Every model/dimensionality-reduction is in some sense compression, isn't it? We are taking a problem space with more parameters and detail and reducing it to a solution space with fewer parameters and possibly less detail (depending on whether the solution is exact or not which maps onto the compression being lossy or not). If I take the set of all points that are equidistant from a given point and line, that is an infinite set. But I can compress it down to three real numbers if I know that set can be represented y = ax^2 + bx +c, and noone goes "quadratic equations are just compression". A lot of people try to generate milage out of the idea of the compression being _lossy_ also. That's an intellectual dead end in the same way. Lots of useful models are a lossy compression.
- inimino 3y agoCompression and prediction are equivalent, and both measure some function of intelligence and knowledge. Current LLMs are heavy on the knowledge side, being trained on ~all the text.
- ComplexSystems 3y agoHow should we unify these things? Most of it seems basically the same - we maximize log-likelihoods instead of calling it minimizing surprisal - but it's clearly the same thing. Are there any other ways to integrate these things together?