3 ms·
It explicitly says in the article that the reason they capped it at ~1B parameters is because that’s the current limit of what they can achieve and hit their la
by atty 4y ago
It explicitly says in the article that the reason they capped it at ~1B parameters is because that’s the current limit of what they can achieve and hit their latency requirements? It has nothing to do with not being interested in scaling it further, as far as I’m aware.
- akomtu 4y agoImo, the current "the bigger the better" trend is limiting further progress. Intelligence finds the simplest model that explains all data, e.g. a small set of equations or rules, while the today's wannabe-AI models are trying to remember all the data in a fuzzy lookup table.
- atty 4y agoI think there’s two parts to this. The first is that there’s a minimum number of parameters required to do an arbitrary task. For any sufficiently complex task (image recognition, large language models, etc), it’s not clear how to find that lower bound. And the bound probably depends on the model chosen. But we do know that the more parameters you add, the more complex of a function you can learn. On the other hand, for many reasons (energy, training/inference deployment complexity, latency, even some sense of model elegance I suppose), we don’t want to massively increase the number of parameters unnecessarily. But again, I don’t think we have great methods to estimate the “ideal” minimum number of parameters for a model to achieve its goals. And what we keep finding is that if you increase the number of parameters, and you increase the training corpus, the model gets more accurate, more impressive, and it’s not stopping. So while I definitely agree that size for size’s sake is wasteful, I also don’t think we necessarily even know how to define “wasteful” for things like large language models right now.