2 ms·
What is the likelihood of overfitting the smaller models? It’s not obvious what the criteria and hyperparams are that prevent that. If there’s no overfitting
by binarymax 3y ago
What is the likelihood of overfitting the smaller models? It’s not obvious what the criteria and hyperparams are that prevent that.
If there’s no overfitting and the results get reproduced then this is a very promising find.
- minimaxir 3y agoTo elaborate on my comment on the other thread, a workaround to overfitting is to train on so much distinct data that the model can't overfit. Newer large datasets optimize for diversity, as the AI industry is slowing coming to the realization that better data is much more important than large amount of bad data for training LLMs. The writeup of SlimPajama, a heavily-deduped dataset, is a good starting point: https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama https://www.cerebras.net/blog/slimpajama-a-627b-token-cleane...
- zmmmmm 3y agothis is sort of the bigger picture observation I have about LLMs in general: that they seem to be massively over-parameterised, which in a classical statistical world would suggest they will readily over fit and it will be hard to get them to generalise beyond regurgitating their training material. Yet this is not the behaviour we see. So in some sense the "magic" we have seen flourish is that LLMs have managed to devise training mechanisms that do let them learn without over fitting in the context of a large excess of parameters. Part of it may be the sheer volume of training material but I suspect even with that you would get straight up regurgitation because it is so hard to stop the input data being degenerate. So other factors about how the training works must be contributing to this. Of course, this is all my semi-layman speculation about this and I'm curious what actually knowledgeable people think.