4 ms·
But...they did that, with Llama 2, and apparently did get better results, at least up to a point. My big WTF is, if you're feeding the same amount of data thro
by cornel_io 3y ago
But...they did that, with Llama 2, and apparently did get better results, at least up to a point.
My big WTF is, if you're feeding the same amount of data through all of them, then to use the same amount of compute for the smaller models you need to run multiple epochs (with the same data). One thing that's always bothered me a bit about "foundation model" LLM training is that it sounds like traditionally they essentially just run a single epoch, and with stochastic gradient descent that's certainly leaving something on the table, probably a lot (and also introduces a lot of path-dependence on the order in which data is presented, what with the cosine learning rate rules).
I really want to know what would happen if, like in smaller models where we can actually get there (convnets for ImageNet classification, e.g.), we ran enough epochs on each of these models to hit the point where validation loss started increasing even as test loss decreased. It seems like we're always squarely in the realm where they're still both decreasing, so everything is severely undertrained, even given the available datasets. It's easy to come up with "laws" for that regime, but they mean nothing other than that we don't have enough compute to properly handle the data.
Big takeaway: if the results from this article are legit, it would suggest that we should really be looking at even smaller models, wouldn't it? And actually be training them to the "risk overtraining" point?
- futurisold 3y agoYou might want to give a read to "Scaling Data-Constrained Language Models" [1]. They basically generalized the Chinchilla scaling law by investigating behavior on multi-epoch runs. [1] https://arxiv.org/abs/2305.16264 https://arxiv.org/abs/2305.16264