4 ms·
> According to NVidia, LLM sizes have been increasing 10X per year for the last few years. Clearly this cannot continue, as the training costs will exceed all
by datpiff 4y ago
> According to NVidia, LLM sizes have been increasing 10X per year for the last few years.
Clearly this cannot continue, as the training costs will exceed all the compute capacity in existence.
The other limit is training data, eventually you run out of cheap sources.
- Ambolia 4y agoThey may also run out of data if they already consumed most of the internet. Or start producing so much of the internet's content that LLMs start consuming what they write in a closed loop.
- warkdarrior 4y agoSure, but they do not need to grow infinitely more, they just need to grow sufficiently to perform better than 90% of humans.
- saiojd 4y agoI would not be so sure about compute capacity? Neural network architectures are still in their infancy, it is very likely that more efficient approaches exist.
- datpiff 4y agoSure, but the small & efficient networks so far have been pretty poor. Edge AI was hugely hyped a few years ago and it petered out. All of the big tech companies seems to be chasing the biggest networks with the largest training sets, that seems to be the direction for now. There could be huge efficiency savings within the implementations but when training costs are already this high it seems naive to think the low-hanging fruit is still there.
- saiojd 4y agoI wouldn't be so confident. Flash attention is a recent, significant improvement to training times and sure looks low-hanging now. I was not familiar with Edge AI, interesting concept. I feel like improving the efficiency of very large models is much more likely. The recent successes will lead to an even larger influx of $ in the short term --- this is an existential threat to Google, after all. We will see where things go!
- YeGoblynQueenne 4y agoWe've had neural nets since 1943. The architectures are not "in their infancy", new architectures have been developing for decades, even entire neural net paradigms (feed-forward nets, recurrent nets, recursive nets, etc etc.). Their scale has also been increasing ever since Hinton and friends rediscovered backprop in the '80s. Neural nets are positively ancient at this point, not "in their infancy"! I don't know why people just keep repeating this complete fantasy as if it were true. Where does it originate from, I wonder? I suspect someone said something like that on social media, their post went viral, and now all of the internet is reverberating with this thing. It's a meme, yes?
- saiojd 4y ago> Where does it originate from, I wonder? I suspect someone said something like that on social media, their post went viral, and now all of the internet is reverberating with this thing. It's a meme, yes? I don't have social media outside of HN. It comes from a few observations: 1) Large models are still improving with increased parameter counts (we do not know where the ceiling is yet; it could be low but it could also be high). 2) Most current architectures train by using all model parameters to produce an output, which is vastly inefficient. While it is not clear how to improve on this in the general case yet, in the simpler problem of NERFs, sidestepping this issue has led to a ~100x improvement in training time. 3) https://mingukkang.github.io/GigaGAN/ https://mingukkang.github.io/GigaGAN/ very recently increased the parameter count of StyleGAN by selecting parameters dynamically at runtime. They improved on previous results by a very, very large margin, at somewhat comparable training times. I stand by my claim: "neural network architectures are still in their infancy, it is very likely that more efficient approaches exist". I am not claiming that AI will become sentient or anything crazy and do not understand why you are associating my point of view with other people. I just said that it is likely that a novel technology will continue to improve (has this it ever NOT been the case for any new technology?).
- YeGoblynQueenne 4y ago>> It comes from a few observations: Well, if you want to know whether neural nets are in their "infancy" you shouldn't make "observations", you should read the literature. It goes back many years. Go to the primary sources, why try to guess and risk guessing wrong, as here? >> I stand by my claim: "neural network architectures are still in their infancy, it is very likely that more efficient approaches exist". Half of your "claim" is incorrect. Don't just double down on it! There's really nothing to "claim" here, the "infancy" or not of neural nets is not a matter of claiming or guessing. Either you know what it is, or you don't.