3 ms·
As an expert in the field: this is exactly right. LLMs are trained to do whole book prediction, at training time we throw in whole books at the time. It's only
by 317070 8mo ago
As an expert in the field: this is exactly right.
LLMs are trained to do whole book prediction, at training time we throw in whole books at the time. It's only when sampling we do one or a few tokens at the time.
- justinator 8mo agowhere do you get these books? honking intensifies WHERE DO YOU GET THESE BOOKS?!
- tasuki 8mo agoThe local library.
- fc417fc802 8mo agoCan anyone even say what a book really is at the end of the day? It's such an abstract concept. /s
- deleted 8mo ago[deleted]
- benterix 8mo agoWe do things, but it doesn't feel right
- TuringTest 8mo agoIsn't that the same as compressing the whole book, in a special differential format that compares how the text looks from any given point before and after?
- 317070 8mo agoThere are many ways to model how the model works in simpler terms. Next-word prediction is useful to characterize how you do inference with the model. Maximizing mutual information, compressing, gradient descent, ... are all useful characterisations of the training process. But as stated above, next token prediction is a misleading frame for the training process. While the sampling is indeed happening 1 token at a time, due to the training process, much more is going on in the latent space where the model has its internal stream of information.
- margalabargala 8mo agoEverything is the same as everything else. It's all just hydrogen and time mixed together.