9 ms·
The "language models don't really understand anything" corner is getting smaller and smaller. In the last few months we've seen pretty definitive evidence that
by maxwells-daemon 5y ago
The "language models don't really understand anything" corner is getting smaller and smaller. In the last few months we've seen pretty definitive evidence that transformers can recombine concepts ([1], [2]) and do simple logical inference using contextual information ([3], "make the score font color visible"). I see no reason that this technology couldn't smoothly scale into human-level intelligence, yet lots of people seem to think it'll require a step change or is impossible.
That being said, robust systematic generalization is still a hard problem. But "achieve symbol grounding through tons of multimodal data" is looking more and more like the answer.
[1] https://openai.com/blog/dall-e/ https://openai.com/blog/dall-e/
[2] https://distill.pub/2021/multimodal-neurons/ https://distill.pub/2021/multimodal-neurons/
[3] https://openai.com/blog/openai-codex/ https://openai.com/blog/openai-codex/
- karmasimida 5y ago> The "language models don't really understand anything" This is still true. By all account, human doesn't need to read 159GB of Python code to write Python, or we simply can't. But it doesn't necessarily indicate language models aren't useful.
- maxwells-daemon 5y agoI would argue humans ingest a lot more than 159GB before they can write code. Most of it isn't Python, and humans currently transfer knowledge a lot more efficiently than NNs, but I suspect that'll change as incorporating more varied data sources becomes feasible.
- pnt12 5y agoWe generalize pretty well. One could say: "it took you 20 years to learn python!", but actually I learned python, Java, c#... Software engineering, machine learning... How to play guitar, how to cook.. .How to speak Portuguese, how to speak English... And thousands and thousands of different things which build on each other. You can give a programmer a few kb of code in a new language and that will give him a small grasp of how it works.
- hackinthebochs 5y agoConsidering the sum total of data and computation that goes in to creating an intelligent human mind, including the forces of natural selection in creating our innate structure and dispositions, it's not obvious that any conclusions can be drawn from the fact that so much data and compute goes into training these models.
- nightski 5y agoHas this transfer of knowledge from one domain to another really been demonstrated by these models/learning processes? I know transfer learning is a thing (I have a couple books on my shelf on it). But it seems far from what you are describing.
- visarga 5y agoDALL-E + CLIP models show a deep understanding of the relation between images and text.
- sbierwagen 5y agoThe AlphaZero algorithm swapped between board games pretty easily. OpenAI could also have been gesturing at this when they named the GPT paper "Language Models are Few-Shot Learners".
- talor_a 5y agothey mention in the demo video that the inspiration for codex came from GPT-3 users training it to respond to queries with code samples. I saw some pretty impressive demos of the original model creating SQL queries from plain questions. I'm not sure if that counts as switching domains, but it's something?
- simsla 5y agoThe problem with this (very popular) argument is that you can't give a CS course to a baby and expect them to get at programming. By the time we see our first line of code, most of us have seen a ridiculous amount of data. We've been trained in problem solving, logical reasoning, maths, natural language processing, ... Hell, we've been trained as pattern matchers since we've been born. By my account, humans actually need a large amount of training data. It might be the knowledge federation and generalisation that we're good at, but I don't think we're a clear winner in data efficiency.
- randallsquared 5y agoTaking 11Mbps [1] as the raw uncompressed incoming data, and assuming 16 hours of waking environment consumption on average (likely high for children), a 13yo has taken in less than 400 TB of information (I used 11 * 60 * 60 * 16 * 365 * 13 / 8.) That's... surprisingly low. [1] https://www.britannica.com/science/information-theory/Physiology https://www.britannica.com/science/information-theory/Physio...
- naresh_xai 5y agoAre we still limiting to visual cues and not the auditory,smell,taste,touch data which we get exposed to?
- FeepingCreature 5y agoVisual input is so dense it's basically not worth tracking the other senses from a data rate pov.
- jdonaldson 5y agoI think intelligence as defined as "mapping inputs into goal states" is pretty well handled by models, and the models may be able to pick and choose states that are sufficient for achieving the goals. However, the intelligence that's created by language models is very schizophrenic, and the human-level reflective intelligence that it displays is at best a bit of Frankenstein's monster (an agglomeration of utterances from other people that it uses to form sentences that form opinions of itself or its world). I think that modeling will help us learn more about human intelligence, but we're going to have to do a lot better than just training models blindly on huge amounts of text.
- visarga 5y agoMaybe we're also >50% Frankenstein monsters, an agglomeration of utterances from other people.
- maxwells-daemon 5y agoAs an add-on to this: I'd encourage anyone interested in this debate to read Rich Sutton's "The Bitter Lesson" (http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html). At every point in time, the best systems we can build today will be ones leveraging lots of domain-specific information. But the systems that will continue to be useful in five years will always be the ones freely that scale with increased parallel compute and data, which grow much faster than domain-specific knowledge. Learning systems with the ability to use context to develop domain-specific knowledge "on their own" are the only way to ride the wave of this computational bounty.
- pchiusano 5y agohttps://rodneybrooks.com/a-better-lesson/ https://rodneybrooks.com/a-better-lesson/ is an interesting retort to the Sutton post.
- 6gvONxR4sf7o 5y ago> The "language models don't really understand anything" corner is getting smaller and smaller. In my mind, understanding a thing means you can justify an answer. Like a student showing their work and being able to defend it. An answer with a proof understands the answer with respect to the proof it provides. E.g. to understand an answer with regards to first order logic, it'll have to be able to defend a logical deduction of that answer. These models still can't justify their answers very well, so I'd say they're accurate but only understand with respect to a fairly dumb proof system (e.g. they can select relevant passages or just appeal to overall accuracy statistics). They're still far from being able to justify answers in the various ways we do, which I'd say means that by definition that they still don't understand with regards to the "proof systems" that we understand things with regards to. Maybe the next step will require increasingly interesting justification systems.
- joshjdr 5y agoI found it on Stack Overflow!
- maxwells-daemon 5y agoLook at the "math test" video. Given the question: "Jane has 9 balloons. 6 are green and the rest are blue. How many balloons are blue?" The model outputs: "jane_balloons = 9; green_balloons = 6; blue_balloons = jane_balloons - green_balloons; print(blue_balloons)" That seems like a good justification of a (very simple) step-by-step reasoning process!
- wizzwizz4 5y agoExcept I could do that with a few regex substitutions, which would not be reasoning. The “intelligence” is in the templates provided by the training data. (Extracting that is impressive, but not that impressive.)
- riku_iki 5y agochances are high that something similar was in training set, and model approximated it.
- Voloskaya 5y agoThe definition of "understanding" behaves just like the definition of "intelligence": The threshold to qualify gets pushed by as much as the technology progresses, so that nothing we create is ever intelligent and nothing ever understands.
- bufferoverflow 5y agoIt probably can scale, but we're nowhere near the computational power we need to even recreate the brain. And don't forget, our brain took a billion years to evolve. A typical brain has 80-90 billion neurons and 125 trillion synapses. That's a big freaking network to train. Hopefully we can figure out how to train parts of it and then assemble something very smart.
- jacquesm 5y agoTakes on average 2.5 decades to train it.
- mattkrause 5y agoThat's just from the most recent checkpoint :-) If you were to build it "from scratch" you'd also need to include the millions of years of (distributed) evolution required to get that particular kid to that point. Tony Zador has some interesting thoughts about that, including"A critique of pure learning", here: https://www.nature.com/articles/s41467-019-11786-6 https://www.nature.com/articles/s41467-019-11786-6)
- lstmemery 5y agoI have to disagree with you here. In the Codex paper[1], they have two datasets that Codex got correct about 3% of the time. These are interview and code competition questions. From the paper: "Indeed, a strong student who completes an introductory computer science course is expected to be able to solve a larger fraction of problems than Codex-12B." This suggests to me that Codex really doesn't understand anything about the language beyond syntax. I have no doubt that future systems will improve on this benchmark, but they will likely take advantage of the AST and could use unit tests in a RL-like reward function. [1] https://arxiv.org/abs/2107.03374 https://arxiv.org/abs/2107.03374
- nmca 5y ago12B, though. What about 1.2T?
- lstmemery 5y agoYou need to scale the amount of data to take advantage of the increase in parameters. I'm not sure where we would find another 100 GitHubs worth of data.
- ruuda 5y ago> but they will likely take advantage of the AST In the end, a more general approach with more compute, always wins over applying domain knowledge like taking advantage of the AST. This is called “the bitter lesson”. http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- kevinqi 5y ago"the bitter lesson" is a very interesting, thank you! However, I wonder if AST vs. text analysis is fully comparable to the examples given in the post. Applying human concepts for chess, go, image processing, etc. failed over statistical methods, but I don't think AST vs. text is exactly the same argument. IMO, using an AST is simply a more accurate representation of a program and doesn't necessarily imply an attempt to bring in human intuition/concepts.
- fpgaminer 5y ago> "language models don't really understand anything" I have a sneaking suspicion that, if blinded, the crowd of people saying variations of that quote would also identify the vast majority of human speech as regurgitated ideas as well. > I see no reason that this technology couldn't smoothly scale into human-level intelligence Yup, the OpenAI scaling paper makes this abundantly clear. There is currently no end in sight for the size that we can scale GPT to. We can literally just throw compute at the problem and GPT will get smarter. That's never been seen before in ML. Last time I ran the calculations I estimated that, everything else being equal, we'd reach GPT-human in 20 years (GPT with similar parameter scale as a human brain). That's everything else being equal. It is more than likely that in the next twenty years innovation will make GPT and the platforms we use to train and run models like it more efficient. And the truly terrifying thing is that, to me, GPT-3 has about the intelligence of a bug. Yet it's a bug who's whole existence is human language. It doesn't have to dedicate brain power to spatial awareness, navigation, its body, handling sensory input, etc. GPT-human will be an intelligence with the size of a human brain, but who's sole purpose is understanding human language. And it's been to every library to read every book ever written. In every language. Whatever failings GPT may have at that point, it will be more than capable of compensating for in sheer parameter count, and leaning on the ability to combine ideas across the _entire_ human corpus. All available through an API.
- raducu 5y agoSorry for my limited knowledge of GPT, but isn't it limited by the training data set, like all other models?
- 8eye 5y agoi am exactly where you are, it’s not a matter of if, merely a matter of when
- andreyk 5y ago"I see no reason that this technology couldn't smoothly scale into human-level intelligence, yet lots of people seem to think it'll require a step change or is impossible." I am a big fan of LMs and am not in the don't really understand crowd, but here are a couple of reasons: 1. Large language models such as GPT or Codex still have several major architectural limitations. They lack the ability to make use of long-term memory, since they have a fairly limited amount of info they can take as input; GPT 3 is great at short stories, but can't go beyond that, and it's hard to prime it with a lot of information as you would eg a new employee. There is some work on this, but afaik not very much and it's very much unsolved. 2. Large language models have only gotten this good by ingesting massive amounts of data and scaling up compute. Yet, this growth comes with diminish returns for every order of magnitude. So it just not being to scale either the data or the compute needs sufficiently (with existing hardware architectures) is a very plausible reason. 3. Large language models 'have it easy' because they only deal with one modality (text). Humans intelligence on the other hand is multimodal - we can process vision inputs, sound, touch, etc. sound, etc. simultaneously and share concepts between these modalities. And we likewise output motor commands that result in motion, text. So far it's not too obvious how to achieve this - OpenAI took a step with DALL-E, but that was by just mining a massive amount of image-text pairs, and it's not obvious this is easy for other modalities, in particular for motor control. 4. Human-level intelligence is often framed as having system 1 (reactive output) and system 2 (longer term reasoning not in response to immediate stimuli) - this is not at all present in language models. 5. related to above two, at least some of human intelligence is derived from reinforcement learning (optimizing a policy that is multi-step with a delayed reward). This is much harder than the plain self-supervised learning of LMs. And probably there are a bunch more like these. So while I do think these sorts of models represent a lot of progress, there are many reasons to be doubtful that just 'scale it up' will work to get much further.
- nxmnxm99 5y agoHumans (technologists) in particular are awful at extrapolating. Transformers being able to combine rudimentary, defined "concepts" and scaling that into human intelligence makes about as much sense as extrapolating a XOR gate or an if-else statement "scaling" to human intelligence. I'm of the "human beings are much more than big linear algebra functions slapped on top of a large processor" crowd.
- marto1 5y agoTo play devils's advocate: well, it might just very well turn out that what most humans are currently engaged in can very well turn out to be reducible to "big linear algebra functions slapped on top of a large processor". And then the part about "being just like humans" will be the marketing gravy train that funds the operation.
- deleted 5y ago[deleted]