4 ms·
Can anyone give some color on to what extent advancements in AI are limited by the availability of compute, versus the availability of data? I was under the im
by greatwave1 4y ago
Can anyone give some color on to what extent advancements in AI are limited by the availability of compute, versus the availability of data?
I was under the impression that the size and quality of the training dataset had a much bigger impact on performance versus the sophistication of the model, but I could be mistaken.
- jacobn 4y agoBoth matter, and returns fall off as you go further in one but not the other. The Chinchilla paper[0] established a simple scaling law for Large Language Models: model size and training tokens should grow at the same pace. Compute is then proportional to the product of model size & data quantity. That said, quality of data also matters a lot - OpenAI has had human labelers produce the data for their Reinforcement Learning from Human Feedback (RLHF), which has probably had a disproportionate impact on the success of ChatGPT compared to previous models, but that data is probably O(1%) of what they trained on. At this point I'm guessing OpenAI are limited by both data & compute. Rumor has it they're training the "next big thing" now and it won't finish until December. If they had more compute they could presumably finish sooner, and if they had more data they would presumably let it train longer. [0] https://arxiv.org/abs/2203.15556 https://arxiv.org/abs/2203.15556
- pixl97 4y agoAlso at this point, very few of the biggest players are going to tell us anything about which matters most and those fine tuning numbers can represent a huge strategic advantage. Forcing your competitors to spend billions in hardware and time can put you far ahead of them quickly, at least at our current rate.
- wsgeorge 4y agoAFAIK it's still an active area of research, and evidence from Meta AI [0] suggests that size and quality of data can let smaller (not necessarily less sophisticated) models do amazing things. But a lot of the advancements we're seeing right now are the result of more sophisticated models [1], and one person is doing some interesting work [2] around achieving transformer-level performance with other architectures. So it's not completely settled if more data is the answer. But it has a significant impact. [0] https://ai.facebook.com/blog/large-language-model-llama-meta-ai/ https://ai.facebook.com/blog/large-language-model-llama-meta... [1] https://en.wikipedia.org/wiki/Transformer_(machine_learning_model) https://en.wikipedia.org/wiki/Transformer_(machine_learning_... [2] https://github.com/BlinkDL/RWKV-LM https://github.com/BlinkDL/RWKV-LM
- gsatic 4y agoIt's like a calf being born. It gets up and starts walking. Pretty amazing. Mesmerises everyone. The model contains everything it needs to know, to walk. But it's not going to dance nor is it capable of working out how to dance. Babies do something completely different. They can't walk when born. Their model is to blunder about and work things out, building the model up thro a can I do this - can I do that - why not etc. Its only through this doing learning happens. We have calf ai right now..you ask the calf what do you want to learn next or what are you curious about and you get to see how dumb it is.