5 ms·
It's not a bathtub curve. Your low-level and "high"-level tasks are the same thing: Probabilistic text generation. It's not reasoning about your code, nor abou
by ADeerAppeared 2y ago
It's not a bathtub curve. Your low-level and "high"-level tasks are the same thing: Probabilistic text generation.
It's not reasoning about your code, nor about the explanation it gives you.
> AI has no concept of "these four parts have to be closely connected, building a whole".
AI can't think. It doesn't create an internal model of the problem given, it just guesses. It fails at all these "middle" tasks because they require abstract reasoning to be correct.
- bubblyworld 2y agoNot going to comment on the thinking part, because who knows what that means, but there's evidence that transformers do in fact learn predictive models of their input space. There's a cool blog post on this here: https://www.neelnanda.io/mechanistic-interpretability/othello https://www.neelnanda.io/mechanistic-interpretability/othell...
- ADeerAppeared 2y agoI should clarify, "of the problem given" refers to the problem given in a prompt. As you note, transformer (and indeed, most ML) models do create a "world model". They're useful for 'specific' intelligence tasks. The problem for general tasks lies in their inability to create specific models. To stick with the board game example: The model can't handle differently shaped boards, or changes to the rules. I could ask a human and a chess-trained AI system to, for a given chess board state and piece, what places that piece can move to. Both have their model of chess. But if I then ask, "With the rule change that the pawn can always move two spaces", the AI cannot update their model. Where for the human this would be trivial. The human can substitue in new logic rules, the AI cannot. And that is very core of what's required for generalized logic and "thinking" in the way most tasks require it. What's so troublesome about current generative AI is that it's trained to be extremely general (within the domain of text generation), so their internal models aren't all that good. Ask an LLM the chess problem above and you might even get a good answer out, but it doesn't generalize to all such chess problems, especially not more complex ones.
- adroniser 2y agoIf you're going to suggest something you think an LLM can't do I think at the very least as a show of good faith you should try it out. I've lost count of the number of times people have told me LLMs can't do shit that they very evidently can.
- ADeerAppeared 2y agoI explicitly say that LLMs could do it in my response. As a show of good faith you should try reading the entire comment. Yes, I'm using simple examples to demonstrate a particular difference, because using "real" examples makes getting the point across a lot harder. You're also just wrong. I did in fact test, and both GPT 3.5 Turbo and 4o failed. Not only with the rule change, but with the mere task of providing possible moves. I only included the admission that they may succeed as a matter of due diligence, in that I cannot conclusively rule out they can't get the right answer because of the randomization and API-specific pre-prompting involved. > "For chess board r1bk3r/p2pBpNp/n4n2/1p1NP2P/6P1/3P4/P1P1K3/q5b1 (FEN notation), what are the available moves for pawn B5"
- adroniser 2y agoI did read your entire comment, and that is what prompted my response, because from my perspective your entire premise was based on LLMs failing at simple examples, and yet despite admitting you thought there was a chance an LLM would succeed at your example, it didn't seem you'd bothered to check. The argument you are making is based on the fact that the example is simple. If the example were not simple, you would not be able to use it to dismiss LLMs. I am not surprised that GPT 3.5 and 4o failed, they are both terrible models. GPT4-o is multimodal, but it is far buggier than gpt-4. I tried with claude 3.5 sonnet and it got it first try. It also was able to compute the moves when told the rule change.
- bubblyworld 2y agoThe paper on Othello is of course a very limited model, useful because it's simple enough to study and complex enough to have interesting behaviour. But the general takeaway is that this is evidence that large transformers like GPT, which are trained to predict text, are fully capable of developing emergent models of parts of that input space whenever it is convenient for minimising the loss function. In practice this means that GPT may have internal models of the semantics of human dialogue that are sophisticated enough for it to get by in the enormous variety of prediction tasks we throw at it. I agree with you that it's likely these internal models aren't very detailed (for the reason you wrote - they're very general). The linked blog actually talks about this at the end - an OthelloGPT trained to be good at Othello rather than just able to play legal moves ends up with a worse board model. Presumably because it needs to "invest" more in playing better moves. But if you agree with the blog's take then this is just a matter of scale and training. And whether it's possible or not for them to develop models capable of complex tasks like strategy games with shifting rules is certainly not something you (or anyone else for that matter) can say with certainty right now. Edit: I should clarify we're using "model" in two senses here. There's the actual transformer model, but what I and the blog are talking about is specific weights and neurons _inside_ these transformers that learn to predict complex features of the input space (like legal moves and board updates in the case of OthelloGPT). These develop spontaneously during the training process, which is why they are so interesting. And why they are not really analogous to the "ML models" you refer to in your first two paragraphs.
- circuit10 2y agoAI is clearly capable of some level of abstract reasoning, because abstract reasoning is necessary for accurate probabilistic text generation
- Workaccount2 2y agoIt's not "whether or not it thinks" its "whether n-dimensional vector multiplication in an intricate embedding space is thinking or not". Which on the surface is easy to knee-jerk a "no" too, but with a bit more pondering you realize that however the brain thinks must be describable by math, and now you need to carve out what math is "thinking" and what math is "computation". Or just be a duelist and attribute it to a soul or whatever.
- orbillius 2y ago> however the brain thinks must be describable by math Roger Penrose believes that some portion of the work brains are doing is making use of quantum processes. The claim isn't too far-fetched - similar claims have been made about photosynthesis. That doesn't mean it's not possible for a classical computer, running a neural network, to get the same outcome (any more than the observation that birds have feathers means feathers are necessary to flight). But it does mean that it could be that, yes you can describe what the brain is doing with math ... but you can't copy it with computation.
- thornewolf 2y agoit feels self-evident that computation can mimic the brain. as a result, it's difficult to argue this line much further. to say the brain is non-computable is to assert the existence of a soul, in my opinion.
- Hasu 2y agoA lot of things feel self-evident then turn out to be completely wrong. We don't understand the processes in the brain well enough to assert that they are doing computation. Or to assert that they aren't! > say the brain is non-computable is to assert the existence of a soul, in my opinion I don't believe in souls, but the brain might still be non-computable. There are more than two possibilities. If it is the case that brains are doing something computable that is compatible with our Turing machines, we still have no idea what that is or how to recreate it, simulate it, or approximate it. So it's not a very helpful axiom.
- dudus 2y ago> It doesn't create an internal model of the problem given, it just guesses. It's not entirely true. They often use some sort of memory/scratch-pad to keep a context other than previous tokens. This recent exploit lets you see claude's default prompt that have some references to this system. https://youtu.be/AbPTz08oq58?si=7F5Lbbkxg99tr3FP https://youtu.be/AbPTz08oq58?si=7F5Lbbkxg99tr3FP
- naasking 2y ago> It's not reasoning about your code, nor about the explanation it gives you. We don't really know what "reasoning" is. Presumably you think humans reason about code, but humans also only have statistical models of most problems. So if humans only reason probabilistically about problems, which is why they still make mistakes, then the only difference is that AI is just worse at it. That's not an indication it isn't "reasoning".
- bena 2y ago“We don’t know how we do it, so we can’t say this isn’t how we do it” isn’t a valid argument. We may not know exactly how we reason, but we can rule out probabilistic guessing. And even if that is a part of it, we’re capable of far more sophisticated models. We can recurse and hold links. We can also make intuitive leaps that aren’t quite built on probability.
- naasking 2y ago> “We don’t know how we do it, so we can’t say this isn’t how we do it” isn’t a valid argument Yes it is, assuming we don't know of any specific things that "this" literally can't do but that we can. Which we currently don't, we merely have suspicions. > We may not know exactly how we reason, but we can rule out probabilistic guessing. No we can't. > even if that is a part of it, we’re capable of far more sophisticated models. Yes, but that would be a difference of degree not of kind. This is what scaling proponents have been saying, eg. that scaling does not appear to have a limit. > We can also make intuitive leaps that aren’t quite built on probability. I don't think we have evidence of that. "Intuitive leap" could just be a link generated from sampling some random variable.
- bena 2y agoBut it’s not replicating results that a human would give you. Since it’s not giving the same type of results, then it’s not doing the same thing. If anything, LLMs have definitively ruled out probabilistic guessing as the model for human intelligence. Even now, you’re trying to force LLMs onto human intelligence. Insisting it is despite it not delivering the results. And I’m sure you believe if we just fire up another few million gpus, we’d get there. But we’ll just get wrong faster. LLMs don’t produce new, they just remix old
- PoignardAzur 2y ago> AI can't think. It doesn't create an internal model of the problem given, it just guesses. These "AI can't think" comments pop up on every single thread about AI and they're incredibly tiresome. They never bring anything to the discussion except reminding us how inherently limited these AIs are or whatever. Someone else already replied with the OthelloGPT counter-example that shows that, yes, they do have an internal model. To which you reply that the internal model doesn't count as thinking or abstract reasoning or something, and... like, what even is the point of bringing that up every discussion? These assertions never come with empirical predictions anyway. GP's comment was interesting because it pointed at a specific area of what LLMs are bad at. A thousandth comment saying "LLMs can't think or do abstract things (except in all the cases where they can but those aren't really thinking)" doesn't bring any new info.