3 ms·
> No one knows for certain how language models that are 10x or even 100x larger than current state-of-the-art ones will perform Read the Deepmind Chinchilla pa
by learndeeply 4y ago
> No one knows for certain how language models that are 10x or even 100x larger than current state-of-the-art ones will perform
Read the Deepmind Chinchilla paper, it answers this
- cs702 4y agoIIRC, no, it doesn't. It redefines language-model scaling as a function of both training dataset size as well as number of parameters, as opposed to only the latter. Also, all predictions extrapolated from past experimental results are hypothetical, subject to future experimental verification. We don't yet have a body of theory good enough to trust those extrapolations.
- kordlessagain 4y agoLinear improvements with logarithm scaling of parameters would make sense, maybe.