3 ms·
Can you explain a bit more on the topic of what happens after the base model?
by criticalfault 1y ago
Can you explain a bit more on the topic of what happens after the base model?
- dcre 1y agoThe base model is a pure next token predictor. It just continues whatever prompt you give it — if you ask it a question, it might just keep elaborating the question. To turn these models into something that can actually chat (and more recently, that can do things like tool calls) they do a second phase of training, including reinforcement learning, which teaches the model to maximize some kind of reward signal meant to represent good answers of various kinds. This reward signal applies at the level of the whole response (or possibly parts of the response) so it is not predicting the most likely next token. I don’t know in an absolute sense how much this ends up changing the base model weights, and it’s surprisingly hard to find discussions of this, I guess because the state of the art is quite secret. But it’s clear that RL is important for getting the models to become useful. This is a reasonable explanation, though as a non-expert I can’t vouch for the formal parts: https://www.harysdalvi.com/blog/llms-dont-predict-next-word/ https://www.harysdalvi.com/blog/llms-dont-predict-next-word/ There are other posttraining techniques that are not strictly speaking RL (again, not an expert) but it sounds to me like they are still not teaching straightforward next token prediction in the way people mean when they say LLMs can’t do X because they’re merely predicting the most likely next token based on the training corpus.
- criticalfault 1y agoThanks for the explanation and the link. I learned something today. I'm definitely not an expert, but to me, RL and other techniques looks like a guide or a constraint on the still 'next token prediction' concept. What I do not get is - is this all about training? Or is this about inference. In any case, this is still an eye opener and I need to study this a bit more. When talking inference, models from huggingface are composed of what then? Because they can do angentic stuff, no?