4 ms·
I'm curious on your point 1, and I tend to disagree. The number of parameters in these large language models is increasing faster than Moore's law. Currently yo
by mcbuilder 4y ago
I'm curious on your point 1, and I tend to disagree. The number of parameters in these large language models is increasing faster than Moore's law. Currently you need a server full of GPUs just to run inference on a PaLM model. How do you see the size shrinking so drastically? Hardware is improving on important factors like power consumption, but inference hardware needs to scale with the size of the models. Don't get me wrong, it's likely that PaLM itself can run on 2032 phone, but the real advances will be in even more scaled up models.
The future of AI will be in the data center for a long time to come. Maybe after some point the models cease to scale up and that point will be where the model would even overfit on the amount of data we can possibly give it e.g. the entire internet. The PaLM authors allude to this in their conclusion
- stult 4y agoThere was just an article from deepmind on HN about this topic the other day[1], but basically IIRC it argues that all of the LLMs are horrendously compute inefficient, which means there’s a ton of room to improve them. So those models will be optimized over time just as the consumer hardware will be improved until eventually one day the two trends will converge. It’s just a question of when that will happen. [1] https://news.ycombinator.com/item?id=30987885 https://news.ycombinator.com/item?id=30987885
- axg11 4y agoCurrent LLMs are very compute inefficient. I also think retrieval transformers can bring a few orders of magnitude in efficiency improvements. Combined with architectural improvements I think we can get there.