9 ms·
> Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore. Dataset size is no
by fpgaminer 3y ago
> Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore.
Dataset size is not relevant to predicting the loss threshold of LLMs. You can keep pushing loss down by using the same sized dataset, but increasingly larger models.
Or augment the dataset using RLHF, which provides an "infinite" dataset to train LLMs on. Limited by the capabilities of the scoring model which, of course, you can scale the scoring model infinitely so again the limit isn't dataset size but training compute.
- midland_trucker 3y ago> Dataset size is not relevant to predicting the loss threshold of LLMs. You can keep pushing loss down by using the same sized dataset, but increasingly larger models. Deepmind and others would disagree with you! No-one really knows in actual fact. [1] https://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-training https://www.deepmind.com/publications/an-empirical-analysis-...
- fpgaminer 3y agoI don't recall the Chinchilla paper disputing my point. They establish "training-compute optimal" scaling laws, but none of their findings suggest that loss hits any kind of asymptote.
- midland_trucker 3y agoPerhaps we're talking past each other, is "loss threshold" a specific term in LLM literature? Merely pointing out that the debate as to whether we are compute or data limited (OP) has not concluded at all; There are lots of compelling theories on relationship between the two.