5 ms·
> The human brain could be viewed as a token predictor No it really couldn't, because "generating and updating a 'mental model' of the environment." is as diff
by usrbinbash 3y ago
> The human brain could be viewed as a token predictor
No it really couldn't, because "generating and updating a 'mental model' of the environment." is as different from predicting the next token in a sequence, as a bees dance is from a structured human language.
The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs that we make sense of not as abstract data points, but as experiences in a world of which we are part of. We also have a pre-existing model summarizing our experience in the world, including their degradation, our agency in that world, and our intentionality in that world.
The simple fact that we don't just complete streams, but do so with goals, both immediate and long term, and fit our actions into these goals, in itself already shows how far a humans mental modeling is from the linear action of a language model.
- FeepingCreature 3y agoBut the human mental model is purely internal. For that matter, there is strong evidence that LLMs generate mental models internally. [1] Our interface to motor actions is not dissimilar to a token predictor. > The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs So just like multimodal language models, for instance GPT-4? > as experiences in a world of which we are part of. > The simple fact that we don't just complete streams, but do so with goals, both immediate and long term, and fit our actions into these goals Unfalsifiable! GPT-4 can talk about its experiences all day long. What's more, GPT-4 can act agentic if prompted correctly. [2] How do you qualify a "real goal"? [1] https://www.neelnanda.io/mechanistic-interpretability/othello https://www.neelnanda.io/mechanistic-interpretability/othell... [2] https://github.com/hwchase17/langchain https://github.com/hwchase17/langchain
- usrbinbash 3y ago> For that matter, there is strong evidence that LLMs generate mental models internally. Limited models, such as those representing the state of a game that it was trained to do: Yes. This is how we hope deep learning systems work in general. But I am not talking about limited models. I am talking about ad-hoc models, built from ingesting the context and semantic meaning of a string of tokens, that can simulate reality and allows drawing logical conclusions from it. In regard to my example given elsewhere in this HN thread: I know that Mike exits the elevator first because I build a mental model of what the tokens in the question represent. I can draw conclusions from that model, including new conclusions whos token-representation would be unlikely in the LLMs model, which doesn't explain anything about reality, but explains how tokens are usually ordered in the training set.
- FeepingCreature 3y agoThe relevant keyword you want is "zero-shot learning". (EDIT: Correction; "in-context learning". Sorry for that.) LLMs can pick up patterns from the context window purely at evaluation time using dynamic reinforcement learning. (This is one of those capabilities models seem to just pick up naturally at sufficient scale.) Those patterns are ephemeral and not persisted to memory, which I agree makes LLMs less general than humans, but that seems a weak objection to hang a fundamental difference in kind on. edit: Correction: I can't find a source for my claim that the model specifically picks up reinforcement learning across its context as the algo that it uses to do ICL. I could have sworn I read that somewhere. Will edit a source in if I find it. edit: Though I did find this very cool paper https://arxiv.org/abs/2210.05675 https://arxiv.org/abs/2210.05675 that shows that it's specifically training on language that makes LLMs try to work out abstract rules for in-context learning. edit: https://arxiv.org/abs/2303.07971 https://arxiv.org/abs/2303.07971 isn't the paper I meant, since it only came out recently, but it has a good index of related literature and does a very clear analysis of ICL, demonstrating that models don't just learn rules at runtime but learn "extract structure from context and complete the pattern" as a composable meta-rule. edit: I think I was thinking of https://arxiv.org/abs/2212.10559 https://arxiv.org/abs/2212.10559 , which asserts that ICL acts equivalent to gradient descent. > In regard to my example given elsewhere in this HN thread: I know that Mike exits the elevator first because I build a mental model of what the tokens in the question represent. I can draw conclusions from that model, including new conclusions whos token-representation would be unlikely in the LLMs model, which doesn't explain anything about reality, but explains how tokens are usually ordered in the training set. I mean. Nobody has unmediated access to reality. The LLM doesn't, but neither do you. In the hypothetical, the token in your brain that represents "Mike" is ultimately built from photons hitting your retina, which is not a fundamentally different thing from text tokens. Text tokens are "more abstracted", sure, but every model a general intelligence builds is abstraction based on circumstantial evidence. Doesn't matter if it's human or LLM, we spend our lives in Plato's cave all the same.
- usrbinbash 3y ago> In the hypothetical, the token in your brain that represents "Mike" Mike isn't represented by a token. "Mike" is a word I interpret into an abstract meaning in an ad-hoc created, and later updated or discarded model of a situation in which exist only the elevator, some abstract structure around it, and the laws of physics as I know them from knowledge and experience. > built from photons hitting your retina, which is not a fundamentally different thing from text tokens. The difference is not in how sensory input is gathered. The difference is in what that input represents. For the LLM the token represents...the token. That's it. There is nothing else. The token exists for its own sake, and has no information other than itself. It isn't something from which an abstract concept is built, it IS the concept. As a consequence, an language model doesn't understand whether statements are false or nonsensical. It can say that a sequence is statistically less likely than another one, but that's it. "Jenny leaves first" is less likely than "Mike leaves first". But "Jenny leaves first" is probably more likely than "Mario stands on the Moon", which is more likely than "catfood dog parachute chimney cloud" which is more likely than "blob garglsnarp foobar tchoo tchoo", which in turn is probably more likely than "fdsba254hj m562534%($&)5623%$ 6zn 5)&/(6z3m z6%3w zhbu2563n z56". To someone reaching the conclusion that Mike left the elevator first by drawing that conclusion from an abstract representation of the world, all these statements are equally wrong. To a language model, they are just points along a statistical gradient. So in a language models world a wrong statement can still somehow be "less wrong" than another wrong statement. --- Bear in mind when I say all this, I don't mean to say (and I think I made that clear elsewhere in the thread) that this mimickry of reasoning isn't useful. It is, tremendously so. But I think it's valueable to research and understand the difference in mimicking reason by learning how tokens form reasonable sequences, and actual reasoning from abstracting the world into models that we can draw conclusions from. Not in the least because I believe that this will be a key element in developing things closer to AGIs than the tools we have now.