3 ms·
I’m guessing one is data. The limit would be once you’ve trained a LLM on all public (or even private) data. Sure you can still make some improvements or try to
by maxlamb 4y ago
I’m guessing one is data. The limit would be once you’ve trained a LLM on all public (or even private) data. Sure you can still make some improvements or try to find some additional private data but still, a fundamental limit has been reached.
- macrolime 4y agoIs it even feasible any time soon to train an LLM on all of YouTube?
- danielbln 4y agoNapkin math, assuming around 156 million hours of video on all of Youtube: 156 million hours of YouTube videos 9,000 words/hour 6 characters/word (including space) First, let's find out the total number of characters: 9,000 words/hour \* 6 characters/word = 54,000 characters/hour Now, let's calculate the total number of characters for 156 million hours of YouTube videos: 54,000 characters/hour \* 156,000,000 hours = 8,424,000,000,000 characters Since 1 character is typically 1 byte, we can convert this to gigabytes: 8,424,000,000,000 bytes / (1024 \* 1024 \* 1024) ≈ 7,842.11 GB So, 8TB of text? Seems doable.
- macrolime 4y agoI mean the actual video, that's much bigger With a vision transformer each token may be around 16x16 pixels. I found an example where they use images of resolution 224x224 for training a vision transformer so if we go with that that 256 pixels per token and 50176 pixels per image, so 196 tokens per frame, 24 frames per second, that's 4704 tokens per second or 16934400 token / hour. In total we're at 2.6x10^15 tokens. GPT-3 was trained on 5x10^11 tokens, so YouTube done this way would be around four orders of magnitude more tokens that GPT-3 was trained on. GPT-3 was undertrained by 1-2 orders of magnitude, so the compute required to trained a model on YouTube would then be around 6 orders of magnitude higher than what was used to train GPT-3, so about one million times more. I did a linear regression on the training costs from cerebras(1) and came up with the formula (1901.67366*X)-197902.72715 where X is number of tokens in billions. Plugging in 5x10^15 tokens we get a training cost of 5 billion dollars. I guess a lot of optimizations could be done that would decrease the cost, so maybe its doable in a few years. 1. https://cirrascale.com/cerebras.php https://cirrascale.com/cerebras.php
- FrojoS 4y agoGood point. But isn't the next logical step to allow these systems to collect real world data on their own? And also, potentially even more dangerous, act in the real world and try out things, and fail, to further its learning.
- pixl97 4y agoThis is likely exactly what we will do... which is very questionable when you may have an unaligned paperclip maximizer hidden in there.
- Espressosaurus 4y agoI'm wondering what happens once LLMs are generating large portions of the internet. What then? It's poisoning its own well at that point.
- great_psy 4y agoI think that’s actually really not a limit. We don’t teach babies by throwing lots of data at them, instead we teach them by giving them useful data. The loop I see is: - train on a lot of existing data - run out of useful data - people use ai and give feedback (we’re here) - perform reinforcement learning on the data collected Loop over the last two steps. There is already more than enough data available, it’s just not nicely labeled to say if it’s high quality or should be discarded. Those last two steps will implicitly do the labeling.
- pixl97 4y agoReality throws a lot of unfiltered data at a baby. Now I guess you can consider things like gravity and pain useful data because of the consequences of violating them. But it's this data that grounds the baby in the world it exists in.