5 ms·
> it is because I understand the stochastic-parrot argument and think it is erroneous. Okay then, what exactly about it is erroneous? Because stochastically so
by usrbinbash 3y ago
> it is because I understand the stochastic-parrot argument and think it is erroneous.
Okay then, what exactly about it is erroneous? Because stochastically sorting the set M of known tokens by likelyhood of being the next, is literally what LLMs do.
- FeepingCreature 3y agoThere's a class of statements that can be either interpreted precisely, at which point the claim they make is clearly true but trivial, or interpreted expansively, at which point the claim is significant but no longer clearly true. This is one of those: yes, technically LLMs are token predictors, but technically any nondeterministic Turing machine is a token predictor. The human brain could be viewed as a token predictor [1]. The interesting question is how it comes up with its predictions, and on this the phrase offers no insight at all. [1] https://en.wikipedia.org/wiki/Predictive_coding https://en.wikipedia.org/wiki/Predictive_coding
- usrbinbash 3y ago> The human brain could be viewed as a token predictor No it really couldn't, because "generating and updating a 'mental model' of the environment." is as different from predicting the next token in a sequence, as a bees dance is from a structured human language. The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs that we make sense of not as abstract data points, but as experiences in a world of which we are part of. We also have a pre-existing model summarizing our experience in the world, including their degradation, our agency in that world, and our intentionality in that world. The simple fact that we don't just complete streams, but do so with goals, both immediate and long term, and fit our actions into these goals, in itself already shows how far a humans mental modeling is from the linear action of a language model.
- FeepingCreature 3y agoBut the human mental model is purely internal. For that matter, there is strong evidence that LLMs generate mental models internally. [1] Our interface to motor actions is not dissimilar to a token predictor. > The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs So just like multimodal language models, for instance GPT-4? > as experiences in a world of which we are part of. > The simple fact that we don't just complete streams, but do so with goals, both immediate and long term, and fit our actions into these goals Unfalsifiable! GPT-4 can talk about its experiences all day long. What's more, GPT-4 can act agentic if prompted correctly. [2] How do you qualify a "real goal"? [1] https://www.neelnanda.io/mechanistic-interpretability/othello https://www.neelnanda.io/mechanistic-interpretability/othell... [2] https://github.com/hwchase17/langchain https://github.com/hwchase17/langchain
- usrbinbash 3y ago> For that matter, there is strong evidence that LLMs generate mental models internally. Limited models, such as those representing the state of a game that it was trained to do: Yes. This is how we hope deep learning systems work in general. But I am not talking about limited models. I am talking about ad-hoc models, built from ingesting the context and semantic meaning of a string of tokens, that can simulate reality and allows drawing logical conclusions from it. In regard to my example given elsewhere in this HN thread: I know that Mike exits the elevator first because I build a mental model of what the tokens in the question represent. I can draw conclusions from that model, including new conclusions whos token-representation would be unlikely in the LLMs model, which doesn't explain anything about reality, but explains how tokens are usually ordered in the training set.
- FeepingCreature 3y agoThe relevant keyword you want is "zero-shot learning". (EDIT: Correction; "in-context learning". Sorry for that.) LLMs can pick up patterns from the context window purely at evaluation time using dynamic reinforcement learning. (This is one of those capabilities models seem to just pick up naturally at sufficient scale.) Those patterns are ephemeral and not persisted to memory, which I agree makes LLMs less general than humans, but that seems a weak objection to hang a fundamental difference in kind on. edit: Correction: I can't find a source for my claim that the model specifically picks up reinforcement learning across its context as the algo that it uses to do ICL. I could have sworn I read that somewhere. Will edit a source in if I find it. edit: Though I did find this very cool paper https://arxiv.org/abs/2210.05675 https://arxiv.org/abs/2210.05675 that shows that it's specifically training on language that makes LLMs try to work out abstract rules for in-context learning. edit: https://arxiv.org/abs/2303.07971 https://arxiv.org/abs/2303.07971 isn't the paper I meant, since it only came out recently, but it has a good index of related literature and does a very clear analysis of ICL, demonstrating that models don't just learn rules at runtime but learn "extract structure from context and complete the pattern" as a composable meta-rule. edit: I think I was thinking of https://arxiv.org/abs/2212.10559 https://arxiv.org/abs/2212.10559 , which asserts that ICL acts equivalent to gradient descent. > In regard to my example given elsewhere in this HN thread: I know that Mike exits the elevator first because I build a mental model of what the tokens in the question represent. I can draw conclusions from that model, including new conclusions whos token-representation would be unlikely in the LLMs model, which doesn't explain anything about reality, but explains how tokens are usually ordered in the training set. I mean. Nobody has unmediated access to reality. The LLM doesn't, but neither do you. In the hypothetical, the token in your brain that represents "Mike" is ultimately built from photons hitting your retina, which is not a fundamentally different thing from text tokens. Text tokens are "more abstracted", sure, but every model a general intelligence builds is abstraction based on circumstantial evidence. Doesn't matter if it's human or LLM, we spend our lives in Plato's cave all the same.