3 ms·
Techniques like distilling show that you can in fact get similar performance with 60% fewer parameters. Yes it depends on making the large model first, but that
by tensor 3y ago
Techniques like distilling show that you can in fact get similar performance with 60% fewer parameters. Yes it depends on making the large model first, but that only implies that somehow the large model is perhaps more useful in initial learning. And if that's the case, then it would be rather surprising if we couldn't identify why and create new training techniques that go directly to the smaller equivalent model.
So in this way you've already lost your bet, as we already know that it's possible to create vastly smaller models with similar performance.
- whimsicalism 3y ago60% is not vastly smaller in the way that I meant - we are talking about differences of three orders of magnitude in this article and that is what I am skeptical of. Nevertheless, I have not actually seen a model exhibiting high-level LM capabilities (of the GPT-4 level) that was distilled. Just because a model can match in perplexity with fewer parameters is not the same as matching the performance as subjectively measured by humanity interacting with the model.
- klyrs 3y agoEarly days, yet. If we could cut the size by 60% every few years, we'd be looking at a situation akin to Moore's law where a performant AI might fit in my grandkid's pocket.
- fnordpiglet 3y agoThat doesn’t mean that a much larger with high quality training data and new training techniques won’t be more powerful than a smaller model. We have incredible efficiencies in modern computers in addition to immense resources. They both are needed for the best performance. I suspect it’s the same for LLM. Superior data, superior techniques, and superior scale will yield superior models.