4 ms·
The confusion arises from the question begging an arbitrary distinction of a token in isolation from all the state in the model and its progression in response
by _wire_ 8mo ago
The confusion arises from the question begging an arbitrary distinction of a token in isolation from all the state in the model and its progression in response to a prompt.
They work because while the process is generating a token at a time, each token has a location in an N-dimensional matrix of overlayed state networks for all the tokens in the training data, the tokens in the given prompt, and the sequence of tokens emitted so far for this prompt.
As an analogy, an image on the screen is emitted a pixel at a time, but each pixel's state is coded as part of a network in a matrix that includes all the residual state from the point of image capture.
And just like an image on your screen the computer has no "ideas" about the contents of the presentation, but other subsystems may use mathematical approaches to selecting and categorizing images, filtering, etc.
The common regard that AI is thinking is purely a matter of appearances, and idiomatic terminology.
As to why we tend to become troubled by the resemblance of AI behavior to thought or creativity, but we are not at all troubled by how entire worlds exist within our TV sets is a matter of surprise and conditioning to the medium.
- sichengo 8mo agoYes, I do admit that I simplified the mechanism in the article, but my question is why the scale of next-token prediction yields reasoning-like behavior.