11 ms·
This is not a response to the Chomsky piece. The main argument advanced by Chomsky et al is that LLMs are neither AGIs nor are they precursors to what we might
by raisin_churn 4y ago
This is not a response to the Chomsky piece. The main argument advanced by Chomsky et al is that LLMs are neither AGIs nor are they precursors to what we might consider AGIs, because, among other reasons, LLMs "learn" differently to how humans do, and that difference comes with strict limitations on the upper bounds of what LLMs can achieve. I'm certainly no expert on linguistics or AI/ML, so I don't know about all that, but this blog post avoids engaging with that claim, and opts instead for ad hominem.
- scotty79 4y ago> LLMs "learn" differently to how humans do Do they? Personally I can't rule out that of LLM model was trained on all of the language a single human heard/read and produced it wouldn't be able to create next utterance that might be indistinguishable from what that human says.
- mjburgess 4y agoYes, that's granted. The issue is that indistinguishable isnt good enough. This is the core problem with this schematised (and i think, pseudoscientific) computer science approach to intelligence. Output isnt intelligent. So, for any given output, it could have been created by system A or system B, whose properties could be radically different. It matters why, eg., we get "I hate the rain!" as output. If system-A says it because it: cares, hates, muses, imagines, prefers, intends... then that's radically different than if B does so because, "it's combining a weather API with some internet chat history".
- scotty79 4y ago> Yes, that's granted. The issue is that indistinguishable isnt good enough. It starts to remind me of "Yes! But it doesn't have a soul!"
- mjburgess 4y agoIf a digital thermometer reads 100C, connected to a black box, are we thereby required to believe that there's boiling water inside the box? Science doesn't deal with the "indistinguishable". We cannot, on earth, simply distinguish between whether we go around the sun, or the sun goes around the earth. Does the solar system have a soul? The world exists, and it has properties, and those are independent of how dumb apes happen to be and what we are in a position to "distinguish" or otherwise. A system generating text is acting as-if its having its intelligence measured. Each sentence we take to be a symptom of its: having a theory of the enviornment, having something to say about it, having some intention, etc. When I say, "I don't like what you're wearing!" that sentence itself isnt somehow "intelligent". It is only a valid measure of my caring, preferring, speaking, intending, thinking... because that is why i said it. A shredder which happened to assemble those words is likewise not intelligent. This is basic science: measurements arent objects; and measurements have validity criteria which is, at least, the causal properties of the system give rise to those measures. In the case of ChatGPT no relevant properties give rise to its ouptut. Its sentences are not caused by any intelligence, and aren't valid measures of it. There is no boiling water. Your digital thermometer is broken.
- genman 4y agoActually we can determine if the Sun goes around the Earth or the other way around - if we can create an better, more accurate model that have larger predictive power then we can assume this model to be more likely to be correct. As I understand, this was one initially of the main issues with the new model proposed by Copernicus - it was not more accurate initially.
- woopsn 4y agoI agree it's a red herring to focus on output and interactive behavior when discussing this. If a "shadow prompt" told chatGPT that it writes at a 3rd grade level, we wouldn't argue as much over how smart the bot is. If it omitted the friendly/helpful/deferential assistant stuff, we'd also argue about it less. Bing's initial defensiveness and aggression made it seem even stupider than the mistakes it was making. They're honing in on better prompts and other configuration that will make the bot seem smarter. It seems smarter to say "I can't answer that question" than to confidently say something untruthful. But the underlying computational program (GPT trained on the internet) is the same. If we judge the program's intelligence based on its output, it isn't well defined. The same thing looks intelligent or hilariously unintelligent based on the tokens you (an intelligent person) provide it with. Or in other words... Suppose we collect all of the system's "intelligent" outputs and disregard the rest. We throw away a lot, the majority of responses, and the resulting set looks impressively smart. The system appears to demonstrate advanced machine intelligence when restricted to (some?) preimages of this set, even though it acts like a total idiot over other parts of the domain. And it's clear that it takes real knowledge and understanding to solve this boundary problem, so that the calculated image has an "intelligent" shape.
- colechristensen 4y agoI think “indistinguishable” is good enough. As ML generated artifacts become more common seeing flavors of what they generate and common failures will become gradually more obvious. It will keep happening that a new technology will seem very impressive and then after a while the cracks will appear and we’ll all have our sort of internal turing test that separates human from machine.
- raisin_churn 4y agoI don't know, that's Chomsky's claim. If I, again a non-expert, had to take a position on it, it seems far more likely to be true than not. Humans have access to a wide variety of non-language stimuli, and demonstrate signs of intelligence well before they have any functional mastery of language. Even after I "mastered" language, I developed lots of skills that haven't the faintest relation to language, like riding a bike. I'm sure ChatGPT can produce a textual explanation of riding a bike, but neither ChatGPT nor a human who doesn't know how to ride a bike can convert that textual explanation into the act of riding an actual bike. But a human, unlike ChatGPT, could, given a bike, learn to ride it by trying to ride it.
- PaulHoule 4y agoI’d look though at systems like CLIP and Stable Diffusion that are able to map between the language domain and images, as well as music, speech, etc. “Riding a bike” can be seen as a sequence modeling problem too because it is a matter of firing muscle fibers in a certain way and it is a research area to make language-controlled robots that do just that.
- zug_zug 4y agoI guess the idea is that if I described a static process to an AI, like multiplication, as we have it wired up right now, it wouldn't be able to remember what I told it for years. I agree this is true, and that it will be a breakthrough for AI, but it's entirely unclear how far away it is in the time dimension.
- pixl97 4y agoIn humans this is the process of taking short term memory and converting it to long term memory involves a process called consolidation where the structure of the physical brain changes, I guess this would be tantamount to a reweighting of the neural net. It's generally not a one shot thing, especially as the concepts get more complex and have more parts to learn. One of the things humans do is forget a lot of unimportant crap so we're not constantly rewriting our brains. Of course there is a the issue of how do we make sure we're training our AI how to learn multiplication and not feeding it a diet of junk food information/fake news too.
- zorked 4y agoThe simple fact that a LLM is trained on a gigantic corpus of data and humans learn from a relatively tiny number of interactions with other humans shows that they obviously learn differently.
- AstixAndBelix 4y ago>humans learn from a relatively tiny number of interactions with other humans but those interactions are infinitely complex and contain an enormous amount of data
- genman 4y agoAnd children are really stupid until they have been exposed to even larger amount of data.
- HDThoreaun 4y agoOr that brains have more processing power or a subtly different architecture than current LLMs
- AndrewPGameDev 4y agoIf we took all of the data that a human takes in from all of their senses, I'm not sure if humans use less data. Humans take in 10 million bits[1] from their eyes every second. 10,000,000 bits/sec * 60 secs/min * 60 mins/hour * 24 hours/day * 1000 days = 108 terabytes. ChatGPT only used 570 GB of training data, so 2 orders of magnitude less data, and that's only counting the visual data. edit: And that would be for a 3 year old, so comparing ChatGPT's intelligence to a 3 year old shows that ChatGPT comes out favourably. [1]https://www.sciencedaily.com/releases/2006/07/060726180933.htm#:~:text=Summary%3A,similar%20to%20an%20Ethernet%20connection https://www.sciencedaily.com/releases/2006/07/060726180933.h....
- deleted 4y ago[deleted]
- v0idzer0 4y agoThat claim deserves no response. Who promised ChatGPT was AGI? Nobody. It’s a straw man argument.
- pharmakom 4y ago> nor are they precursors to what we might consider AGIs
- Calavar 4y agoNobody claims that ChatGPT is AGI, but plenty claim that LLMs may be the the path to AGI.
- davewritescode 4y agoThis is a claim I hear repeated both in meatspace and online from people who I generally regard as really smart. I agree that AGI is quite a ways off.
- raisin_churn 4y agoLoads of people have intimated that ChatGPT shows we're on the precipice of a breakthrough in that direction. Chomsky says we're not. If Chomsky's argument is a strawman, Aaronson could write a blog post saying that, instead of a blog post stuffing his own strawman. Anyways, I'm summarizing Chomsky's argument as I understood it, you can read it yourself.
- scott_s 4y agoNot any of the people who developed it. But some of the lay-public is under this impression, so it’s an important clarification to make.
- colechristensen 4y agoBlake Lemoine was the dude claiming Google’s LLM had a soul and got fired and all sorts of press attention. It’s not a straw man. Blake was an outlier but also very much on the inside and not just some random wonk. There are plenty of people who are being led to believe that LLMs are much more than fancy statistics.
- fwlr 4y agoThe Chomsky article does argue that LLMs are not AGI or AGI precursors: “These programs have been hailed as the first glimmers on the horizon of artificial general intelligence … That day may come, but its dawn is not yet breaking, contrary to what can be read in hyperbolic headlines and reckoned by injudicious investments”. You correctly point out that the counter-article does not respond to those arguments. But there are many other claims and many other arguments the Chomsky piece makes: “we fear that the most popular and fashionable strain of A.I. — machine learning — will degrade our science and debase our ethics by incorporating into our technology a fundamentally flawed conception of language and knowledge”. The Aaronson article is responding to those claims, even if the response is “I disagree; more details to follow”. Separately, it’s rather audacious to dismiss this response as ad hominem, considering the tone of what it’s responding to. Aaronson’s ad hominems: “the intellectual godfather of an effort that failed for 60 years”, “what Chomsky and his followers are ultimately angry at is reality itself”. Chomsky’s ad hominems: “ineradicable defects”, “lumbering statistical engine”, “stuck in a prehuman or nonhuman phase of cognitive evolution”, “the predictions of machine learning systems will always be superficial and dubious”, “pseudoscience”, “ChatGPT exhibits something like the banality of evil”, “the amorality, faux science and linguistic incompetence of these systems”.
- morelisp 4y agoNone of Chomsky’s statements you quote are ad hominems. They’re not even misidentified ad hominem fallacy fallacies (which is what Aaronson’s are, actually). I’m not sure you know what an ad hominem is.
- fwlr 4y agoI should have known that Chomskyites would be bitterly pedantic. Replace “ad hominem” in my post with “mean stuff”.
- oldgradstudent 4y agoIgnoring your ad hominem attack, ad hominem means attacking the speaker rather than the argument, which is what Aaronson does here, and Chomsky does not.