3 ms·
From GP, i.e. the context for this local part of the thread > Autoregressive LLMs generate tokens one at a time, disputing this is just plain wrong. next-toke
by dijksterhuis 22d ago
From GP, i.e. the context for this local part of the thread
> Autoregressive LLMs generate tokens one at a time, disputing this is just plain wrong.
next-token prediction i.e. the bit built during pre-training.
at no point in your reply to GP did you specify that you were referring to post-training. respectfully, it seems like this one is on you pal :shrug:
> GPT-2 didn't use any reinforcement learning and is often given as a toy example. That release was 2019 and models now go through a various phases of training with different objective functions and optimizers.
yeah. so? the toy example works for pre-training. see above.
- danielmarkbruce 22d agoAll modern LLMs that actually get used go through post-training. The finished product is something which has been through post training. So they are not next token prediction machines.
- dijksterhuis 22d ago> The finished product is something which has been through post training. again, the finished product wasn't what was discussed by GP, and you didn't clarify that you were switching to discussing RL (which is still probabilistic btw)
- danielmarkbruce 22d agoYes, it was. Nobody says a system or product works a certain way and means the system while it's half built. "Bridges drop cars in the water!". Right. You aren't in this field. You are clearly wrong and just can't handle it.
- dijksterhuis 22d ago> Nobody says a system or product works a certain way and means the system while it's half built. "Bridges drop cars in the water!". Right. To understand how an engine works, it's important to understand what a piston does as part of the engine.
- danielmarkbruce 22d agoYou are conflating "half built" with "a piece of a system". The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere else. It's more like pottery - the thing changes. It's not correct to say something is soft and malleable because it once was.
- deleted 22d ago[deleted]
- deleted 22d ago[deleted]
- dijksterhuis 22d ago> The model weights change as the model goes through the training process. Yes. They do. You are absolutely right about that. But the model architecture doesn't change as a result of the training process. A piston doesn't suddenly turn into a digital watch as a result of tuning an engine. Similarly, the transformer part of a GPT model doesn't suddenly turn into something else as a result of optimizing a loss function. --- i've got other stuff to do, so i'm stopping here.
- danielmarkbruce 21d agoNo one is arguing about the architecture of the model. It's the objective function and optimizer.
- doc_ick 21d agoJust skimming through here but I think you have the wrong ideas with llms, I’d recommend Andrew Ngs course (correct me if you’ve already seen it or something similar).
- danielmarkbruce 21d ago
- deleted 21d ago[deleted]