4 ms·
rubs temples It's a large language model. It is not smart or dumb. It models the input it is trained on. It is not figuring anything out. It doesn't know anyth
by ttpphd 3y ago
rubs temples
It's a large language model. It is not smart or dumb. It models the input it is trained on. It is not figuring anything out. It doesn't know anything. It isn't reasoning. It is generating text.
When will the Eliza fever break here?
- Turing_Machine 3y ago> It models the input it is trained on...It doesn't know anything. The same is true of any computer program. That doesn't make software in general not useful.
- newswasboring 3y agoWhat a thing is and what it can do are two different things. I can just as easily say CPUs are just fancy rocks.
- Aperocky 3y agoWhat a thing is defines the limit of what it can do. Right now we know what it is, but people are arguing that it has capability arguably beyond the limit. It's akin to arguing that humans can survive without oxygen, and then coming up with some alternative definition of oxygen, or surviving to validate the statement.
- newswasboring 3y agoYes but your line of reasoning shuts down any discussions of what it is because you have tied it to it's manufacturing process. Because it is trained as a word predictor doesn't mean it's all it is. During manufacturing we don't look at CPUs as processing units we just look at it as hunks of rocks we are carving.
- Aperocky 3y agoCPU are fairly predictable in its capabilities. LLMs are not too different in that regard, but harder to quantify, or maybe due to it being very new.
- newswasboring 3y agoMy point isn't about what it is, it's about having the right kind of discussions. I'm seeing people falling into camps and it's way to early for that. My objection isn't that you are wrong, my objection is to the amount of certainty in your original post.
- Aperocky 3y agoIt won't break, some will even argue pattern matching is itself intelligence, I've heard plenty of that. Though if it is it's very different from us, since I can barely recall what happened in the morning.
- flangola7 3y agoWhat definition of intelligence do we use? At this point I've seen 30 different versions and none of them are consistent.
- Aperocky 3y agoI'd just go by ability to solve novel, general logical problems.
- flangola7 3y ago>novel, general logical problems. What might be a couple examples of that?
- renewiltord 3y agoThe risky thing is we might end up with many humans failing. And would that mean they lack intelligence?
- nicpottier 3y agoCan you formulate a few for us and we can try them out?
- radium3d 3y agoThe reason these GPT large language models stand out is, we "generate text" based on the "input text" we are trained on too. But these can train on orders of magnitude more text than we will ever ingest in our entire lifetimes.
- zzzeek 3y agoAssuming you are responding to the article and not the comments here, i think you should read the article as the author agrees entirely with what you are saying.
- ttpphd 3y agoI am not sure I agree with that assessment. The author describes a large language model as having unconscious thoughts.
- antiterra 3y agoThe author explicitly states that analogies to human thought are inaccurate. The purpose of the analogies is describe a mental model of the kinds of things GPT-4 can do or not do.
- ttpphd 3y agoThen let's say the point of disagreement between me and the author is that I would prefer to abandon the analogy for something that lets me both reason and communicate more clearly, while the author appears to prefer to keep the analogy they have identified as flawed.
- newswasboring 3y ago> while the author appears to prefer to keep the analogy they have identified as flawed. All analogies are flawed, some are useful. /Smug
- baq 3y agoThe author is describing a computer with a very peculiar computation model and lots of information stored in ROM making analogies to how humans solve problems along the way.
- red75prime 3y agoSlaps the roof of the head. This bad boy can generate so much text. Text generation is GPT-4's function. How it performs its function is another question.
- ttpphd 3y agoThe algorithm is defined. It has to be for us to use it. The question is what patterns in the training input are being exploited by the algorithm to generate text modeled after the input.
- Mathnerd314 3y agoWell it is reasoning. I asked ChatGPT to sum two large numbers from random.org, 490277348+718085950, and it got the right answer 1208363298. The numbers have no Google results showing such an addition. So ChatGPT at least has learned addition. I'm sure if someone analyzed the network carefully enough they could probably find the digit add/carry neurons. In contrast, Eliza was keyword-based and had much less capability to generalize to novel input.
- smoldesu 3y ago> So ChatGPT at least has learned addition. Or, OpenAI added a calculator function. Furthermore, I would add: - ChatGPT hasn't "learned" anything outside of what exists in it's model. - If it is an AI-based response, it's still derived from token-based inferencing. - One run of asking ChatGPT something is not enough to prove much of anything. > I'm sure if someone analyzed the network carefully enough they could probably find the digit add/carry neurons. They will find tokens for 'dig' and 'it', as well as 'add', 'car' and 'ry'. They will not find internalized understanding of the concept of math.
- famouswaffles 3y ago>Or, OpenAI added a calculator function. Furthermore, I would add: No lol. The test doesn't have to be mental arithmetic. and accuracy mistakes creep in at large numbers for multiplication. That's not how calculators work i'm sure you know >ChatGPT hasn't "learned" anything outside of what exists in it's model. I'm sorry...and you have? >If it is an AI-based response, it's still derived from token-based inferencing. Um...Ok? Lol >One run of asking ChatGPT something is not enough to prove much of anything. You can run this on gpt-4 as much as you like. the results are the same. It knows addition. >They will find tokens for 'dig' and 'it', as well as 'add', 'car' and 'ry'. They will not find internalized understanding of the concept of math. That's...not how that works lol, You don't probe neurons and see tokens. https://clementneo.com/posts/2023/02/11/we-found-an-neuron https://clementneo.com/posts/2023/02/11/we-found-an-neuron Weights don't store the data they train on like that. They are essentially configuration settings.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- baq 3y agoMarkov chains also generate text and GPT-4 absolutely isn’t a Markov chain judging purely by its output. GPT-4 is useful. ‘It’s just a bunch of equations’ just isn’t a good argument against this tech.
- ftxbro 3y ago> GPT-4 absolutely isn’t a Markov chain I mean it is technically a high order Markov chain, or at least every published GPT-N is so far. For example, GPT-3 is a 2048-order Markov chain over tokens. There may be some confusion because of the difference between the technical definition of Markov chain vs. how low-order Markov chains are usually implemented. Maybe you are used to Markov chains that have explicit transition matrices stored in memory and they are trained by counting the number of times a token appears after every prefix. That does give a Markov chain. But technically Markov chains aren't required to have their transition matrix be explicitly stored in memory, and they aren't required to be trained by simple counting (max likelihood transitions). They can have implicit transition matrices and be trained by gradient descent or whatever and still technically be Markov chains. I saw Karpathy weigh in on this and he said it's a Markov chain in the same way that a computer is a Markov chain. I guess his point is that if you have a high enough order Markov chain, the intuitions and connotations of Markov chains become less useful, and maybe 2048 order is high enough that it crosses that threshold.
- psyklic 3y agoKarpathy discusses viewing GPT as a finite-state Markov chain in this Twitter thread (and especially in the intro to the linked Colab notebook) : https://twitter.com/karpathy/status/1645115622517542913 https://twitter.com/karpathy/status/1645115622517542913
- ftxbro 3y agoThe comment I had in mind was when he said "Yes but in the same way as saying that computers are just a Markov chain." in response to "Is GPT simply a Markov chain?" https://news.ycombinator.com/item?id=35506756 https://news.ycombinator.com/item?id=35506756
- darthrupert 3y agoIt clearly has emergent properties which make it emulate smartness at the very least. A bit like a game which "just renders 60 images per second" is just an image generator but also a simulation of some aspects of reality.
- avindroth 3y agoBy a similar line of reasoning, many people did not expect LLMs to achieve these results. Maybe reductionism isn't too useful when we are dealing with emergent behaviors.
- iamflimflam1 3y agoI used to say this as well, I’ve written blog posts and made videos saying exactly this. But - I would strongly recommend trying GPT4. There’s a very good video that is worth watching as well that may not change your mind, but will certainly make you think: https://youtu.be/qbIk7-JPB2c https://youtu.be/qbIk7-JPB2c The thing to remember is that these models are unbelievably huge and deep. We don’t know what is really happening in the layers or what it has really learned. Thinking that it’s a simple language model that is just predicting the next most likely word is unwise.
- snewman 3y agoI don't think this sort of dismissal adds anything to the conversation. I myself, in the act of typing this comment, am "generating text". I think it's interesting to discuss the capabilities of leading-edge language models like GPT-4, because (a) they are already exhibiting the ability to perform a wide variety of useful tasks, and (b) it's clear that there is still a lot of unrealized potential here. Can you clarify the implications you see here? Are you saying that these LLMs are somehow uninteresting or incapable? That there are limits to what they will be able to accomplish even with further improvements? Or something else?
- ttpphd 3y agoI am not making a comment about the usefulness of LLMs to manipulate text. If anything, my view is that to the extent that LLMs are seen as intelligent it's precisely because they manipulate text in ways humans find useful, not because it follows from some philosophy about intelligence or rationality. Maybe I'm being too opinionated but I think we should stop dressing up explanations of large language models in misleading terminology. I'd prefer instead to talk about the actual technology and reason from there.
- Loeffelmann 3y agoFor something to generate text that is believable on a level like gpt-4 does it needs to have some model of the real world and understand relationships between things. So yes the training goal was to "just predict the next word" same as our training goal was just "reproduce and survive". What emerged out of that is the real important thing. No human can understand what really happens in the billion of calculations done for each token in gpt-4 so how can you claim that there is surely no thought process going on? It can solve some riddles, it can draw pictures and it can reason (to some extend). How is that just generating text to you? In the end this argument doesn't matter because how it was made is irrelevant. What matters is what it is and can do.
- ttpphd 3y agoThe God of the gaps is now accessible via API.
- danaris 3y ago"I don't understand how the things GPT-4 does could be possible without genuine understanding an a model of the world, so it must have those things (despite there being no evidence from its building blocks or the structure with which it is built that it could be capable of those things)."
- newswasboring 3y agoI haven't seen proof that it could be capable of those things so I will boldly assert it's definitely not doing that. Everyone needs to calm down and realize we don't need to have all the answers right now.
- fxj 3y agoAlso it is important to note that GPT shows large miss-alignments. The problem comes from the fact, that it is hard or impossible to give an objective what GPT should be optimized to (nobody knows the truth), so proxys are used. One proxy is that it should make the user happy and the user should give many thumbs up. But this does not mean that it has to give the "correct" answer, which the user himself might not know in the first place. So it invents things because during the reinforcement learning users were happy with these answers. A funny example is the github co-pilot which writes buggy code, because it thinks this is what the user wants. Here is a video about that: https://youtu.be/viJt_DXTfwA?t=1767 https://youtu.be/viJt_DXTfwA?t=1767