4 ms·
I think there’s two parts to this. The first is that there’s a minimum number of parameters required to do an arbitrary task. For any sufficiently complex task
by atty 4y ago
I think there’s two parts to this. The first is that there’s a minimum number of parameters required to do an arbitrary task. For any sufficiently complex task (image recognition, large language models, etc), it’s not clear how to find that lower bound. And the bound probably depends on the model chosen. But we do know that the more parameters you add, the more complex of a function you can learn.
On the other hand, for many reasons (energy, training/inference deployment complexity, latency, even some sense of model elegance I suppose), we don’t want to massively increase the number of parameters unnecessarily. But again, I don’t think we have great methods to estimate the “ideal” minimum number of parameters for a model to achieve its goals. And what we keep finding is that if you increase the number of parameters, and you increase the training corpus, the model gets more accurate, more impressive, and it’s not stopping.
So while I definitely agree that size for size’s sake is wasteful, I also don’t think we necessarily even know how to define “wasteful” for things like large language models right now.