3 ms·
Right, but it leads to too many false conclusions by lay people. User facing LLMs are only trained on next token prediction during initial stages of their trai
by fpgaminer 2y ago
Right, but it leads to too many false conclusions by lay people. User facing LLMs are only trained on next token prediction during initial stages of their training. They have to go through Reinforcement Learning before they become useful to users, and RL training occurs on complete responses, not just token-by-token.
That leads to conclusions elucidated by the very article, that LLMs couldn't possibly plan ahead because they are only trained to predict next tokens. When the opposite conclusion would be more common if it was better understood that they go through RL.
- mentalgear 2y agoWhat? The "article" is from anthropic, so I think they would know what they write about. Also, RL is an additional training process that does not negate that GPT / transformers are left-right autoencoders that are effectively next token predictors. [Why Can't AI Make Its Own Discoveries? — With Yann LeCun] (https://www.youtube.com/watch?v=qvNCVYkHKfg https://www.youtube.com/watch?v=qvNCVYkHKfg)
- pipes 2y agoListening to this today, so far really good. Glad I found it. Thanks.
- TeMPOraL 2y agoYou don't need RL for the conclusion "trained to predict next token => only things one token ahead" to be wrong. After all, the LLM is predicting that next token from something - a context, that's many tokens long. Human text isn't arbitrary and random, there are statistical patterns in our speech, writing, thinking, that span words, sentences, paragraphs - and even for next token prediction, predicting correctly means learning those same patterns. It's not hard to imagine the model generating token N is already thinking about tokens N+1 thru N+100, by virtue of statistical patterns of preceding hundred tokens changing with each subsequent token choice.
- fpgaminer 2y agoTrue. See one of Anthropic's researcher's comment for a great example of that. It's likely that "planning" inherently exists in the raw LLM and RL is just bringing it to the forefront. I just think it's helpful to understand that all of these models people are interacting with were trained with the _explicit_ goal of maximizing the probabilities of responses _as a whole_, not just maximizing probabilities of individual tokens.