4 ms·
They talk specifically of scaling laws investigated in papers like OpenAI's "Scaling Laws for Neural Language Models" (2020)[1] and Deepmind's "An empirical ana
by airgapstopgap 4y ago
They talk specifically of scaling laws investigated in papers like OpenAI's "Scaling Laws for Neural Language Models" (2020)[1] and Deepmind's "An empirical analysis of compute-optimal large language model training" (2022)[2] which showed that, as far as we could see at the time, for any sensibly designed language model nuances of architecture do not matter nearly as much as the sheer scale of the model, amount of compute spent on training it, and volume of training dataset (or, indeed, that they do not matter at all, or even that attempts to add some cleverness via architecture ultimately handicap the model at production scale). These papers have introduced so-called "scaling laws" – predictable, robust relationships between scaling parameters and loss (which is meaningfully related to textual coherence and apparent "understanding" in zero-shot tasks), and for many have heralded the age of throwing compute at the problem without academic tinkering, and possibly the final stretch before human-level text processing. See also Rich Sutton's Bitter Lesson [3].
In this sense it is remarkable when the curve predicted from "scaling laws" gets broken through, transcended, since it suggests there still exists a better Pareto frontier for text prediction. I see how this can appear normal; but the thing is, scaling laws have been well-validated. We are no longer in the early era to expect to reap substantial returns from something as simple as changing training objective.
1. https://arxiv.org/abs/2001.08361 https://arxiv.org/abs/2001.08361
2. https://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-training https://www.deepmind.com/publications/an-empirical-analysis-...
3. http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- 6gvONxR4sf7o 4y agoIt seems like every other paper in the scaling laws lit finds different scaling laws (if you consider the constant factors important).