5 ms·
> The thing with these models is you can't untrain just 3% of the input, you have to start from the beginning You think they don't checkpoint the model before
by Turing_Machine 3y ago
> The thing with these models is you can't untrain just 3% of the input, you have to start from the beginning
You think they don't checkpoint the model before feeding in a new slug of data?
That would seem...unwise.
On the other hand, the history of IT is rife with people doing unwise things, so it could be true.
- supermatt 3y agoAnd if the first "slug of data" contains something you need to remove? If you need to remove something you would need to roll-back to the last checkpoint that didnt include it.
- mynameisvlad 3y agoI mean, sure, but the point of frequent checkpoints is that these issues would be rooted out sooner rather than later. It would have been caught far sooner than “we have to redo everything from scratch”.
- Turing_Machine 3y agoWell, that's why you checkpoint a lot and feed in the new training data a little at a time, rather than dumping in a massive slug all at once. Right?
- supermatt 3y agoNot really. If they are ingesting peoples data to train their model, at what point are they checkpointing? Is data I submit 2 months ago that I would want removed not present in a recent checkpoint or would they have to "unwind" the last 2 months of training to remove it? What about people that would want something removed 6 months prior? What if something in the core model is to be removed? Checkpointing may help, sure, but it isnt going to allow you to remove a single piece of training data without headache, and potential substantial retraining. If that 3% from the GPs comment comment is uniformly distributed throughout the training history, the only way it can be reliably removed is to retrain from scratch.
- Turing_Machine 3y ago> If they are ingesting peoples data to train their model, at what point are they checkpointing? Multiple times per day if they're not incompetent. And I don't think they're incompetent.
- supermatt 3y agoAre you implying that removing the diff between checkpoints achieves the same effect? Ive never heard of this, but I suppose it may be possible. I suppose the "ghost" of the removed weights would also have shaped subsequent training though... Interesting idea...
- visarga 3y ago> You think they don't checkpoint the model before feeding in a new slug of data? They definitely do that otherwise they can't rewind the model when the loss shoots up. Unstable training happens to almost all LLMs, but can be managed by rewinding & skipping a few batches.