3 ms·
> Why do scaling laws work? Strictly speaking, the original paradigm of scaling laws doesn't work any more. The assumption that we could achieve better perform
by mentalgear 5mo ago
> Why do scaling laws work?
Strictly speaking, the original paradigm of scaling laws doesn't work any more. The assumption that we could achieve better performance simply through "vertical scaling" ie infusing models with exponentially more parameters and pre-training data, is no longer the driving force of AI progress.
Instead, the industry has pivoted toward inference-time scaling. Rather than relying solely on a massive, static neural network, modern architectures allocate more compute during the actual generation process, allowing the model to "think" and verify its logic dynamically.
Furthermore, the latest state-of-the-art models are no longer pure LLMs; they are compound neuro-symbolic systems that integrate external tools like REPLs, databases, and structured skill documentation to archive things pure LLM vertical parameter scaling was not able to do.
- yorwba 5mo agoThe "law" part of scaling laws is about predicting validation cross-entropy loss from the training configuration, analogous to physical laws allowing to predict one quantity based on the measurement of another. Most scaling laws take the form of an irreducible error plus additional terms that asymptotically decay to zero. So that there is a wall you can approach but not cross (the irreducible error) is an integrated part of the scaling law paradigm. That it isn't economical to keep increasing model size to squeeze out a few more drops of cross-entropy doesn't mean scaling laws stopped working. Strictly speaking, "Why do scaling laws work?" is a question about the theoretical reasons the asymptotic decay takes the particular mathematical shape that it does.