11 ms·
Does this imply we will run out of data to keep up with larger model sizes? Is there much more data out there than what they’re already using?
by mrfusion 5y ago
Does this imply we will run out of data to keep up with larger model sizes?
Is there much more data out there than what they’re already using?
- adamsmith143 5y agoProbably not an issue just yet, think of how much data is generated by Twitter on a daily basis for example.
- teraflop 5y agoOn the other hand, consider the difficulty of taking massive amounts of data from the modern web and filtering out the subset that was actually generated by humans, rather than previous generations of language models.
- adamsmith143 5y agoDefinitely an interesting future problem. I'm sure OpenAI and others are thinking about it but I don't think these models are ubiquitous enough to have much impact just yet.
- axg11 5y agoSome estimates: - 500M tweets per day - 30 words/tokens per tweet - 40% of all tweets thrown away due to being duplicate/spam/bots = 9B tokens generated per day
- zarzavat 5y agoIf you want to teach your kid to learn English, and they came back to you and said "Dad/mum, I finished reading the entire internet but I still don't understand English fully", would you say "OK son, now go and stare at the Twitter firehouse until you grok perfect English" ? It's clear that these models have orders of magnitude too much data already. It somewhat reminds me of the proposals for larger and larger colliders in the hopes of seeing new physics that is always one collider in the future.
- mijoharas 5y ago> It somewhat reminds me of the proposals for larger and larger colliders in the hopes of seeing new physics that is always one collider in the future. I agree with your main point, but think this analogy isn't an apt one. If you want to see what particles are created at higher energies you kinda need the bigger particle accelerators. (This isn't to say that we shouldn't be investigating lower energy collisions, but at a certain point you do need "bigger colliders" to see new things)
- lostmsu 5y agoI disagree with this take because you grok English not only from the text you read, but also from the context of physical world around you. And that context is enormous: assuming 8000x8000x2 vision with 3 color 1 byte channels at 24fps without compression, you get 3e+17 bytes (300 petabytes) of data along with your reading per year.
- ralfd 5y agoBlind children can learn english fine though. And there are areas highly unmaterial (mathematics) which people still reason about.
- lostmsu 5y agoYou ignored the point. I only brought sight as an example (though, admittedly, it is the largest data inflow).
- hugh-avherald 5y agoHumans have to learn sight while learning speech though. I don't mean knowing that this is a dog, I mean making sense of noisy vision inputs.
- lostmsu 5y agoSo what? There's no indication that one hinders the other much, and may even improve generalization.
- kelseyfrog 5y agoIt implies our models are wrong. Consider that a human adolescence is ~9.46x10^6 minutes and a fast speaking rate is ~200words/minute. That sets an upper bound of 1.9 billion words heard during adolescence. ie: human adults are trained on a corpus of less than 1.9B words. To some extent, more data can offset worse models, but I don't think that's the regieme we're currently in. GPT-3 was trained (on among other languages) 181 billion English words - or about 100 times more words than a human will hear by the time they reach adulthood. How is the human brain able to achieve a higher level of success with 1% of the data? 1. https://github.com/openai/gpt-3/blob/master/dataset_statistics/languages_by_word_count.csv https://github.com/openai/gpt-3/blob/master/dataset_statisti...
- nynx 5y agoYeah, this implies backpropagation is deeply suboptimal.
- kelseyfrog 5y agoThat is certainly a possibility. The other (non-mutually exclusive) implications may also be that human language acquisition benefits from being part of a multi-task model. Or that the problem has been overreduced ie: human language acquisition cannot simply be distilled into a words-in->words-out problem and that vision/hearing are actually integral parts of language acquisition that cannot be left out. Or that model arch still has major improvements to be made and attention is not all you need, for example.
- fpgaminer 5y ago> and that vision/hearing are actually integral parts of language acquisition Deaf-blind authors would beg to differ. But yes, a human brain is exposed to lots of other sensory input, and we know from other research that multi-modal models can learn shared representations that benefit from the knowledge of each domain. In Transformer's favor, at least, they are far closer to tabula rasa than the human brain is and likely have to dedicate a lot of their training time to things that are otherwise "baked" into human brains. For example, humans come pre-packaged with V1 and V2 as part of their visual system, but CNNs and ViTs have to learn those filter packages from scratch. I agree with you though. Human brains are able to take single instances of experiences and build a wealth of understanding from them in ways that even modern Transformer architectures are not yet able.
- replygirl 5y agoThere's a ton of data that can be exponentially more useful, but we'll need networks that can (analogously) be late to work enough times to get fired, or experience heartbreak in succession while misunderstanding why prior heartbreak happened, or hallucinate stray cats when they're walking around the neighborhood at night