15 ms·
A great demonstration of why in-distribution learning is not enough for AGI. The largest GPT-3 model "GPT-3 175B" is trained with 499 Billion tokens (one token
by MAXPOOL 6y ago
A great demonstration of why in-distribution learning is not enough for AGI.
The largest GPT-3 model "GPT-3 175B" is trained with 499 Billion tokens (one token is maybe equal to 4 characters in text).
Human reading/talking/listening equivalent of 200 pages of text per day for 80 years would be just 13GB of raw data or 3B tokens. You could also make an estimate using 39 bits/s as the normal information rate humans can absorb and get the same order of magnitude estimate.
It's not completely wrong to say that we are figuring out how far interpolation from data can go as a learning method. Even the most advanced deep learning is in-distribution learning and system 1 thinking (Kahneman's term). Just run through data until the model can interpolate accurately between data points.
We must figure out how do learn models that allow out-of-distribution learning or these models fail the Turing test after a few questions.
- nxpnsv 6y agoWell, yes - but humans learn from more than just reading. Kids know an awful lot of stuff even before they can read a single word...
- MAXPOOL 6y agoEven in the terms of total human sensory input data rate, humans learn from a very tiny data set. 16 year old human has experienced only about 400 million wakeful seconds.
- so_tired 6y agoSee my other comment > Humans get a hand crafted curriculum inputs, evolved over 1000s of iterations, in a near-optimal language encoding.
- DavidSJ 6y agoAt what bitrate though? At 1 MB/s that’s 400 TB, which is nothing to sneeze at.
- dane-pgp 6y agoTo give a sense of how accurate that estimate is, consider this reference: "In other words, the human body sends 11 million bits per second to the brain for processing, yet the conscious mind seems to be able to process only 50 bits per second." https://www.britannica.com/science/information-theory/Physiology https://www.britannica.com/science/information-theory/Physio...
- DavidSJ 6y agoSo quite accurate then, since 11Mb/s = 1.375MB/s. (The conscious mind, of course, is not where most of our learning from experience happens.)
- pmoriarty 6y agoI am very suspicious of the numbers and statistics quoted in this article, as there are no references cited for their claims, and no indication as to how they arrived at them.
- akiselev 6y agoThat doesn't sound accurate. The human eye's resolution is in the tens if not hundreds of megapixels [1] with high sensitivity (meaning the compression isn't very lossy so the bits of information remain). Even if you assume only a few frames per second that's far more than 11 megabit/s. [1] https://clarkvision.com/articles/eye-resolution.html https://clarkvision.com/articles/eye-resolution.html
- sudosysgen 6y agoHuman vision is more complicated than that, there is some "compression" happening at the eye. But yes, it's almost certainly more than 11 megabits, more likely around a few dozen just for the eyes: https://link.springer.com/chapter/10.1007/978-3-642-04954-5_7 https://link.springer.com/chapter/10.1007/978-3-642-04954-5_...
- browsergap 6y agoInteresting to consider tho...there must be some sort of power threshold. Above which all our questions can be merely "interpolated" from data. Right now 499B tokens is not it. But there must be, I guess, some upper limit, within which all human knowledge expressible through language can be contained and conversed upon using this method. Pretty scare to think about...that at that point, when it's 10^n tokens or whatever, we would be unable to detect if it understood or not. Even more scary, what if our brains are simply above that power limit in their ability (if they do that) to interpolate. And what if we don't really understand anything (but, just like vision, our brains provide us the comforting sensation that we really get it), but simply are in possession of wetware/quantum computers that can interpolate better than we can poke holes in it.
- killerstorm 6y agoThere's no evidence that human intelligence is anything more than associative memory and very elaborate multi-level pattern matching. IMHO any cognitive task can be described as a combination of associative lookups, domain transformation and trial-and-error stateful processing. I don't see any reason to believe there's more to intelligence/consciousness than a combination of such elements/processes. I guess domain transformation is something which is hard to visualize. But it was visualized in style transfer: there are NNs which can decompose a picture into a subject and style and then take the same subject and apply a different style. Same works with text -- you can take a sentence and rewrite it to be in Victorian epistolary style. Or take a story about humans and rephrase it to be a fairy tale with animals. GPT-3 can also do this. This means it's possible to take a sentence and decompose it into different layers, then manipulate layers individually and reassemble.
- browsergap 6y agoCool. I don't like to include consciousness in this reduction, I consider that something sacred. The field that permeates all of us and everything. The thinking part of it, I think there's a chance it could be as you describe, with our brains coming with algorithms (or the ability to form similar algorithms) that do domain transformation, stateful processing, associate lookups, memory. I like to think of the consciousness, as simply being aware of this, and I guess "gently guiding" it through intention. It's what tunes us into our thinking and let's us perceive it. But the thinking could occur without that "supervisor". But perhaps somehow consciousness is the spark which gives us life, because when it leaves the body, the body and brain no longer function, or at least no longer animate. And consciousness can seem to exist within and outside the body. Maybe what we call intelligence, and what we are trying to create with AI is just "how a brain" thinks. That version of AI definitely seems achievable, it's a;sp interesting to think what consciousness is, when it seems that people have been able to perceive things outside their body, while it was dead, and then return to that body upon it being revived, and being able to bring verifiable details back with them...Or a consciousness being able to recall things inbetween incarnations in bodies....Those impressions must involve pattern matching and thinking as well...But there's no substrate it seems. Just the universe itself. Which is pretty bizarre...Makes you wonder why our brains need to be so complex at all, if our souls can think and remember too... Perhaps even more scary and bizarre (and not what I believe) but what if having a soul (as in some consciousness part that exists separate to the body) is actually real but not common. So the strange anecdotes of these occurrences are not more widespread simply because, for whatever reason, most bodies don't have souls...I'm sure it's not like that...but since we're speculating, may as well take this to where it goes, which is, if that's the case then consciousness (which we all have, otherwise the body is anesthetized), and the soul (which it seems some people have, owing to how they have disembodied memories that are factual) might be separate. So if the body has consciousness that's entirely embodied, maybe AI's can have consciousness as well. Then what's the soul? And what is the soul doing when a person's brain and consciousness is doing the thinking? Or if we all have a soul (a disembodied consciousness) why can only very few people seem to have disembodied memories? If we make the right AI substrate for consciousness, will it attract souls, like moths to a light?
- killerstorm 6y agoI don't think it makes sense to compare human learning to GPT-3 learning: it's a fundamentally different process. Human brain doesn't get just tokens, but also other sensory data, particularly, visual. So I don't think that you can conclude that humans learn more efficiency based on just quantity of data. It's also worth noting that GPT-3 is trained to emulate _any_ human writing, not just some human's writing. For an actual Turing test one might fine-tune it on text produced by one particular human, then you might get more accurate results.
- MAXPOOL 6y ago> I don't think that you can conclude that humans learn more efficiency based on just quantity of data. You are correct. My main argument was that in-distribution learning is not enough. You can's fix that problem with more data as many responses to my comment seem to assume. I think out-distribution leaning and small data requirement are connected. If agent can understand the concept separate from the chain leading from sensory inputs, it can understand what it's doing in novel situation even without examples.
- BoiledCabbage 6y ago> You can's fix that problem with more data as many responses to my comment seem to assume. From my reading, you've asserted this to be true, but haven't given any more support to why then the people who have asserted it to be false. GPT-3 isn't a human brain. I want to understand your argument as more than "planes can't fly because they don't flap their wings".
- killerstorm 6y agoElements used in GPT-3 are capable of transformations such as abstraction (e.g. separate the structure of syllogism from concrete nouns) and logic (ReLU can directly implement OR, AND, NOT which is sufficient to do arbitrary logic). We can see that it actually uses abstractions and logic in some cases. E.g. "Bob is a frog. Bob's skin color is ___". Even small GPT-2 models can relate "Bob" to concept "frog" and query "frog" "skin color" attribute. Even basic language modeling requires inference, and GPT-x can inference using transformer blocks. With more layers it can go from inferencing meaning of words to doing inference to solve problems. But the inference it is able to do is limited in scope because of the structure of a language model -- each input token must correspond to one output token. So the model can't take a pause and think about something, it can only think while it produces tokens. Here's an absolutely insane example of embedding symbolic computation into a story which lets GPT-3 to break computation into small steps it can handle. Intermediate results become part of the story: https://twitter.com/kleptid/status/1284069270603866113 https://twitter.com/kleptid/status/1284069270603866113 https://twitter.com/kleptid/status/1284098635689611264 https://twitter.com/kleptid/status/1284098635689611264 So I guess one can make a model which is much better at thinking simply by training in a different way or changing the topology. But the building blocks are good enough.
- phreeza 6y agoIt's not really a fair thing to compare to a single human life though, since humans benefit greatly from the genetic and cultural priors inherent in the entire course of evolution. The structural prior for a transformer is really quite minimal, the equations fit on an index card.
- MAXPOOL 6y agoThat's true. I don't think humans are very good at general intelligence. Our genes give us priors to evade lions and pick berries. Humanity as cultural organism can reprogram us to more abstract tasks, but it takes easily 10-20 years of training and we are relatively inefficient in them compared to our original task. But even then. I don't think current in-distribution learning is enough even if we scale it up. There must be something else. We are at least 3-5 Turing awards away from AGI, I think.
- jfengel 6y agoProbably. But we were more like 10 to 20 Turing awards way just a decade ago. I don't expect progress to be consistent; those remaining awards could take another century. But we made what might be a huge amount of progress in a very short time. Or it could be a complete dead end, as has happened before. And nobody could tell until it was done; even those who were correctly skeptical were really just lucky that there wasn't a rabbit in that hat. Sometimes there's a rabbit.
- phreeza 6y agoIn the general sense, is out-of-distribution learning even possible? The best you are likely going to get is robustness wrt different environments, but the environments have to come from the same Meta-Distribution. Invariant risk minimization is an interesting research area in that direction, there is for example an adversarial formulation that seems quite elegant. I have never tried this, but I can imagine it is tricky to get to converge.
- MAXPOOL 6y ago>In the general sense, is out-of-distribution learning even possible? Yes. Humans can learn to answer questions even when questions (test cases) don't come from the same distribution as the training set. You need to have some amount of analytical capabilities. Analytic as the ability to form distinct and modular concepts from the input and form representations that can be composed together to generate instances that don't exist in the distribution. The low-level example of this ability could be blind source separation.
- so_tired 6y ago> Human reading/talking/listening equivalent of 200 pages of text per day for 80 years would be just 13GB of raw data or 3B tokens I am sympathetic to the "total life time input" argument. But humans get a hand crafted curriculum inputs, evolved over 1000s of iterations, in a near-optimal language encoding. Also, if unsupervised gets us in-dist, and DRL seems to be not-bad in search-out-of-dist.... then we are getting close ? Certainly a x100 scaling of current techniques can get a useful enough machine that makes many human tasks trac-able? (I am not getting into the Turing / AGI / skynet argument)
- codingslave 6y agoThe human brain is not a blank slate, it already has an architecture built in that is custom made to handle certain types of information. Asserting that we are learning from scratch is a ridiculous statement, the learning of the human brain has been occurring for millions of years through evolution.
- tgv 6y agoAnd so is GPT-3.
- lumost 6y agoHuman's also don't read redundant data. How many nearly identical news articles did GPT-3 have to read? How many redundant paragraphs from different wikis, reddit comments etc.