30 ms·
> The finished product is something which has been through post training. again, the finished product wasn't what was discussed by GP, and you didn't clarify t
by dijksterhuis 29d ago
> The finished product is something which has been through post training.
again, the finished product wasn't what was discussed by GP, and you didn't clarify that you were switching to discussing RL (which is still probabilistic btw)
- danielmarkbruce 29d agoYes, it was. Nobody says a system or product works a certain way and means the system while it's half built. "Bridges drop cars in the water!". Right. You aren't in this field. You are clearly wrong and just can't handle it.
- dijksterhuis 29d ago> Nobody says a system or product works a certain way and means the system while it's half built. "Bridges drop cars in the water!". Right. To understand how an engine works, it's important to understand what a piston does as part of the engine.
- danielmarkbruce 29d agoYou are conflating "half built" with "a piece of a system". The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere else. It's more like pottery - the thing changes. It's not correct to say something is soft and malleable because it once was.
- deleted 29d ago[deleted]
- deleted 29d ago[deleted]
- dijksterhuis 29d ago> The model weights change as the model goes through the training process. Yes. They do. You are absolutely right about that. But the model architecture doesn't change as a result of the training process. A piston doesn't suddenly turn into a digital watch as a result of tuning an engine. Similarly, the transformer part of a GPT model doesn't suddenly turn into something else as a result of optimizing a loss function. --- i've got other stuff to do, so i'm stopping here.
- danielmarkbruce 29d agoNo one is arguing about the architecture of the model. It's the objective function and optimizer.
- doc_ick 29d agoJust skimming through here but I think you have the wrong ideas with llms, I’d recommend Andrew Ngs course (correct me if you’ve already seen it or something similar).
- danielmarkbruce 29d agoSo, this is the cause of the problem.... People take an intro to LLMs course, follow happily along, and don't realize there is more to it than the next token prediction. And those courses teach how LLMs were built in 2017-2020 maybe. Then RL got added to the mix. The current models really are very different to the models from then - everything that is now considered "post-training" isn't doing next token prediction.
- deleted 28d ago[deleted]
- doc_ick 28d agoPlease feel free to cite sources then, otherwise I see no relevancy from you.
- Dylan16807 29d agoYou're using the fact the both parts of training affect the same weights to support your argument that they're making the system do something fundamentally different after RL?
- danielmarkbruce 29d agoAssuming you are saying that RL is changing the model from doing one thing to another, yes. RL is changing the nature of the model.