6 ms·
I feel like your observation that this "isn't a complicated question" is leaning on an implicit assumption that ChatGPT is a general AI and not a LLM. It is ju
by gwright 3y ago
I feel like your observation that this "isn't a complicated question" is leaning on an implicit assumption that ChatGPT is a general AI and not a LLM. It is just generating text based on probabilities -- it isn't "reasoning". I might go as far to say that inferences computed by a LLM are of all the same complexity but I don't really know enough about ChatGPT to be confident in that statement.
- nearbuy 3y agoPeople keep repeating that LLMs are "just generating text based on probabilities". That statement doesn't mean anything. I think people who say this are imagining LLMs work something like a statistical model. Maybe it's doing a linear regression or works like a Markov chain. It's not. A single artificial neuron sort of works like that. But that's sort of like saying a single transistor is just an electronically controlled switch, so the only thing computers can do is switching. It's true in some sense that computers are just doing a lot of switching, but it turns out all this switching is Turing-complete. That means computers can theoretically compute anything that's possible to compute given enough time and memory, which includes anything a human could figure out. Similar principle applies to LLMs. Using probabilities is part of what they do, but that doesn't preclude them from using logic and rules of inference.
- vidarh 3y agoIt's worth noting that "just generating text based on probabilities" describe Markov algorithms [1], which are Turing-complete. People overestimate how much it takes to end up with something Turing-complete.[Markov algorithms only generate text with probability 100% or 0% based on whether a certain rule matches or not, so it's even simpler] (A Markov algorithm is distinct from a Markov chain, but as far as I can tell you could emulate a Markov algorithm with a Markov chain with sufficient number of input states, transitions clamped to 0% or 100%, and allowing it to iterate over its own output; with a large enough state machine, iteration, and a mechanism to provide memory it's almost hard not to end up with a Turing machine) [1] https://en.wikipedia.org/wiki/Markov_algorithm https://en.wikipedia.org/wiki/Markov_algorithm
- simonh 3y agoThat's a valid point in that we don't fully understand how LLMs solve some problems and using logic and rules of inference isn't excluded by the architecture, but on the other hand understanding that they are generating probabilistic token sequences is a very powerful and effective way to understanding how to engineer prompts and understand some of their failure modes. If we discard that insight, reasoning about their many failure modes and limitations becomes near impossible. For example we often see people thinking that because an LLM can explain how to do something that therefore it knows how to do it, like arithmetic. That's because if a human can explain how to do something, we know that they can. Yet for an LLM outputting a token sequence for an explanation of something, and outputting a token sequence for solving a problem statement for that problem domain are fundamentally different tasks. We can get round this with very clever prompt engineering to 'force' chain of reasoning behaviour, as this discussion shows, but the reason we have to do that is precisely because the cognitive architecture of these LLMs is fundamentally different from humans. Yet these systems are clearly highly capable, and it is possible to dramatically improve their abilities with clever engineering. I think what this means is that LLMs may be incredibly powerful components or elements of systems that may become far more advanced and sophisticated AIs. However to do that engineering and build dramatically more capable systems, we need to have a clear understanding of how and why LLMs work, what their advantages and limitations are, and how to reason about and work with those features.
- vidarh 3y ago> For example we often see people thinking that because an LLM can explain how to do something that therefore it knows how to do it, like arithmetic. That's because if a human can explain how to do something, we know that they can. I think example shows LLMs to be more like people not less. It's not at all unusual to see humans struggle to do something until you remind them that they know an algorithm for doing so, and nudge them to apply it step by step. Sometimes you even have to prod them through each step. LLMs definitely have missing pieces, such as e.g. a working memory, an ability to continue to learn, and an inner monologue, but I don't think their sometimes poor ability to recall and follow a set of rules is what sets them apart.
- 3y ago
- guenthert 3y ago> That means computers can theoretically compute anything that's possible to compute given enough time and memory, which includes anything a human could figure out. Whoa, that's quite a leap there. Not sure where we (as society) are with our understanding of intuition, but I doubt a million monkeys would recognize that the falling of an apple is caused by the same agent as the orbit of planets.
- nearbuy 3y agoI think you misunderstood. I'm not making any claim there. I'm just defining what Turing-complete means for those who don't already know.
- deleted 3y ago[deleted]