12 ms·
33TB of text data for a 1T-parameter model
- lisper 3y agoIf anyone needed to be convinced that LLMs are not accurate models of the human brain they need look no further than these numbers. The smallest model under discussion requires "more books than are in the Kindle store on Amazon U.S." Humans obviously acquire language proficiency with a lot less input than that.
- flashgordon 3y agoTrue but the machinery (is the brain) that enables humans to acquire/process/understand that knowledge is a lot more complex and went through years of environmental conditioning?
- lisper 3y agoYes, but that environmental conditioning did not involve training on the contents of the Kindle library. And the results of that training are encoded in under a gigabyte. For the record, I'm not saying LLMs are not a huge step forward. They are. But they are not -- and cannot be -- the whole answer.
- GaggiX 3y agoI think the human brain is naturally predisposed to learn a language, an evolutionary bias. That helps tremendously compared to a tabula rasa.
- sgu999 3y agoIf we are going in that direction, I guess the few million years of evolution should also count
- flashgordon 3y agoNo totally. Not an ml guy by long shot. My understanding is llms are foundational models which iiuc are "base" models representing all human knowledge until some point in time. Now I'd argue that you don't have to count all of evolution but just the point from when you were born. I'd say the all the sensory experiences + knowledge input from the day a brain is activated (DOB) would be in the order of PBytes if not more?
- deleted 3y ago[deleted]
- scarmig 3y agoOne argument (that I don't necessarily buy) is that inputs from sensory perception play a major role in building a language model in humans. A human might not read TB of books to learn language, but if you put them in a sensory deprivation tank that only displayed a massive series of Unicode characters it would take forever for us to learn language. If we did the same for an ML model, perhaps those non-language training materials would help. (That said, DL architectures are obviously wildly different from how the human brain works. E.g. backprop is physically impossible.)
- ml_basics 3y agoYep totally agreed. One of the things I'm excited to see going forward is multimodal models (trained e.g. on text + video + audio + images). I'm sure there's a lot more to it than this, but maybe one factor that makes humans a lot more data efficient is the multimodal input we receive. If that's the case, imagine how much better things could get when we train with all the videos, podcasts, radio etc in the world, in addition to all the text out there!
- fatneckbeard 3y agoHelen Keller managed to do pretty well without sight or hearing.
- youssefabdelm 3y agoTouch is pretty important though...
- ttpphd 3y ago[flagged]
- undersuit 3y agoShe was brought up in respect to a human in a sensory deprivation tank. You've made the connection to LLMs.
- mindvirus 3y agoFair point, but accurate versus efficient are different. While humans do acquire language much more efficiently, these models "know" more than a human about most things - in that if you asked it to describe each of the 100 most popular subjects or books it would have no trouble doing so (in a dozen languages no less). So I don't think it's quite an apples to apples comparison.
- lisper 3y ago> these models "know" more than a human about most things So does Wikipedia. But that's not a good model of the human brain either. > I don't think it's quite an apples to apples comparison. Yes. That is exactly my point. Despite the superficially similar I/O behavior, the two systems are very different under the hood.
- fastball 3y agoI don't think anyone has argued that LLMs are effectively digital humans or digital human brains, so seems like you're strawmanning a bit here.
- krainboltgreene 3y agoYou aren't reading enough of HN/Twitter then.
- int_19h 3y agoTwitter is a madhouse, but as for HN, the closest I've seen anyone come to that here is claiming that the difference in underlying structure can/does lead to emergence of the same phenomena.
- stevenhuang 3y agoNot as strong an argument as you think as you're forgetting humans have 5 senses, are embodied, and millions of years of tuned weights from natural evolution codified into our brain structure.
- neatze 3y agoBrain is nothing like attention networks at least in terms of dynamics and individual neuron complexity, my wild ass guess; there is not enough space in DNA to store even 10% of brain connections. What weights are you referring to ?
- BobbyJo 3y agoCluster density, cluster location, eagerness to connect, distribution of cell types, etc. are all coded in DNA and contribute to creating "initial weights". Neuron "weights" need not be stored completely uncompressed. Obviously newborns need to develop and take in stimulus before they "know" anything, so the initial conditions are not sufficient, but they are obviously necessary to make the limited learning useful.
- neatze 3y agoIt always seems naive to compare neural structures and functions to artificial networks, dynamics are simply not there, even single neural would require multi layered network (if not mistaken 7 layer network, and very long training time), it is fascinating how brain self-(re)organizes through out life time from small number of examples here and there. As far my understanding goes to very large degree it is unknown how to model such dynamics, where it would be possible to start with no spiking neurons and evolve effective/stable learning behavior from small number of examples/experiences.
- stevenhuang 3y agoFor all we know it may not even be necessary to model the higher fidelity aspects of our brains for a form of intelligence to emerge. There may be many roads to intelligence. The path biology and evolution took may just be one such path.
- gumballindie 3y agoThis is a hallucination.
- kolinko 3y agoOur brains are come pre-trained for speech from a few million years of evolution. Artifitial NN are trained from scratch. A chimpanzee can listen to humans as much as it wants, and it will still not pick up much of the language.
- jxy 3y agoRight. The nature trained all the nervous systems with genetic algorithms for millennia. Individuals rather fine tuned their own with natural noise and dropoffs and natural feedback. Fine tuning our brains are rather expensive.
- est31 3y agoThere are a few innate reflexes but you lose them quickly and they are pretty basic. Our brains have way way more neurons/parameters than these models and are still not overfitting... the issue lies in the learning method.
- famouswaffles 3y agoIt's not just about reflexes though I would hardly classify reflexes as "lost quickly". That's just not true. It's deeper than that. Our brains are optimized for a lot general human functions. Learning language is one of them. There's a whole section of the brain dedicated to it and other things. There's also a lot of vital biological information encoded in DNA/RNA. We are not even close to starting from scratch.
- deleted 3y ago[deleted]
- mrshadowgoose 3y agoThat argument is utterly unconvincing, as you've conveniently left out the millions of years of "training" and "fine tuning" that has occurred through the evolutionary process.
- nl 3y agoWhich implies the brain has more than a structural bias for language; ie language and/or knowledge is somehow passed genetically. There's some reasons to suspect this is at least partially true, but to what extent is unknown and contraversial.
- marginalia_nu 3y agoIt could be argued based on the extreme difficulty we're having in teaching animals human language, even primates. We can create a semiotic system and get a chimpanzee to communicate through it, but it won't learn English no matter how hard we try.
- nl 3y agoI think that's generally accepted to be that humans have a very large language center in the brain; IE: and inductive bias towards learning language (and I think there is something about our mouth and tongue being able to make specific sounds?) It doesn't say anything about that knowledge being precoded.
- marginalia_nu 3y agoIsn't a language center by itself precoded human toward the acquisition of human language (which has co-evolved to make use of our language center)? I don't understand how it could be any other way.
- nl 3y agoExactly. But it is unclear if there is intrinsic language knowledge somehow encoded too (which is what the OP is implying: https://news.ycombinator.com/item?id=35590091 https://news.ycombinator.com/item?id=35590091)
- Dylan16807 3y agoThat depends on what you mean by "accurate". If you mean a lot of accuracy, that's obvious and doesn't really need argument. And this new fact doesn't change the argument. If you mean a more moderate amount of accuracy, this isn't proof either way. Human brains take in less text but they put a lot more processing into it.
- jacquesm 3y agoTrue, but humans have a 4B element text book downloaded into each cell and some of that stuff encodes for a machine that already has a lot of language proficiency built in from day #1. I've always wondered how much of that genetic code is 'soft' in the sense that it pre-sets certain bits in the brains that it generates. Sort of a boot-strap ROM for a human.
- mupuff1234 3y agoDoes someone claim that they are? But regardless, the architecture of the human brain doesn't have to be the only way to get to AGI (not that LLMs are necessarily the way)
- teruakohatu 3y agoYet we learning from 18 hours of ultra high resolution video per day (x2), along with 18 hours of reinforcement learning, along with 18 hours of audio and 18 hours of nerve stimulus data. This is assuming we are learning nothing during sleep, which probably isn't true. By the time a person is 21 years old, they have been trained on at least 1 petabyte of data. By two years old, about 125TB of data. It makes LLMs look quite good in comparison.
- abetusk 3y agoBy no means rigorous, an estimate for the number of neurons is around 10^10 [0] with an average number of connections estimated at around 8000 [1]. Call 10^10 \approx 2^40 for convenience, and 8000 \approx 2^13, which gives us a 2^53 entropy estimate, or about a petabyte of information as an estimate of what the human brain can store (discounting more exotic theories of memory stored in DNA or some such). [0] https://en.wikipedia.org/wiki/Human_brain#Microanatomy https://en.wikipedia.org/wiki/Human_brain#Microanatomy [1] https://psychology.stackexchange.com/questions/7967/how-many-synapses-in-the-average-human-brain https://psychology.stackexchange.com/questions/7967/how-many...
- edulix 3y agoAdd to that that our neurons are more complex to point neurons used in typical Artificial neural-nets. A single pyramid neuron in the neocortex might be more comparable to a multilayer neural net. https://www.biorxiv.org/content/10.1101/2021.10.25.465651v1.full.pdf https://www.biorxiv.org/content/10.1101/2021.10.25.465651v1....
- pmoriarty 3y agoTheir communication is much more complex too, and neither they nor their connections are static. We don't understand how they work at the subatomic level simply because human understanding of the subatomic world is not complete, but even just at the atomic level a single neuron is massively more complex than anything humans have created. Going up to the molecular level, even that is staggeringly more complex than the incredibly simple abstractions that make up a neural net. Is what happens in the brain at the molecular, atomic, or subatomic levels relevant or necessary to intelligence and consciousness? We just don't know yet, but we do know all of that is far more complex and very different from the simple abstractions that are used for neural nets and LLMs. The back of a napkin calculations in this thread don't even begin to do justice to the tremendous amount of "calculation" or "storage" that happens in the human brain.
- H8crilA 3y agoWhy do they have to be accurate models? They just have to be possible to construct and have to perform. The construction method is irrelevant beyond feasibility. I don't care if my model needs an exabyte of RAM if I can just go and buy that much RAM one day.
- deleted 3y ago[deleted]
- est31 3y agoThey aren't accurate models of the human brain just as airplanes are not flapping their wings and cars don't walk on legs. Our brain's way of achieving intelligence is not the only one out there, even if LLMs aren't there yet. Yes, we use now magnitudes more oil than we used to 120 years ago. But back then oil was used for cooking and lamps, while now it is used for so many more things. Same goes for data. It is the new oil :).
- fnordpiglet 3y agoI’d note that the models don’t just acquire language facility, but the ability to use language with alacrity about almost any subject. There’s no human alive that is remotely as capable, and most I’ve met hallucinate more.
- marginalia_nu 3y agoWhile I agree with the conclusion, I don't think your argument actually supports it. Humans for the most part don't acquire language through books until they're mostly already fluent in their native language. These LLMs are also trained on arbitrary books. When acquiring new languages, we use educational materials that are specifically created to facilitate an understanding of language. Not just arbitrary books in random order.
- deleted 3y ago[deleted]
- tyingq 3y agoI'm curious how they keep LLM generated text from turning into future training input, and creating a loop that probably isn't good for quality. Or is that not a problem?
- VHRanger 3y agoThat's a problem for machine learning systems in general. For instance, a recommendation system's output effects what users see, so it effects what they click on. The next training set is statistically dependent on the input of the previous. There are strategies to deal with it, but in the case of a LLM it seems difficult apart from downweighing everything after 2023 that isn't from a vetted source.
- comboy 3y agoOpenAI was talking about some kind of steganography so that previous outputs can be excluded from the training data. Not sure what's the progress on that.
- javajosh 3y agoJust ignore anything that ends with an "In conclusion" paragraph.
- teaearlgraycold 3y agoI will begin all of my writing with “as a large language model” to effectively opt out of all training sets.
- FartyMcFarter 3y agoAs a large language model, that sounds really clever. Let the cat and mouse games begin.
- navjordj 3y agoI would be very suprised if this hasn't already been implemented since GPT-3/3.5
- akomtu 3y agoAfter those 33TB get tokenized and the tokens are encoded with basic frequency coding, I bet that much less than 1TB is left. This makes me think that LLMs are essentially LZW compression on steroids where text is indexed by meaning. It allows to query the data by meaning, but the catch is that every query needs to run a sort of matrix transform on the entire dataset (even though it's reoresented in a compressed form on the LLM weights). Edit: Still, the idea of mapping symbols and words to a many-dimensional space of meanings is a great insight into how mind works. In that space, symbols with similar meaning appear next to each other, and a thought looks like a smooth but intricate shape that separates all symbols into the "insiders" and "outsiders" that, in practice, divide symbols into true/false, good/bad and so on. Such smooth intricate shapes appear in the frequency domain as a bunch of rational numbers, and that's the boundary of what a mind can imagine.
- h2odragon 3y agomarkov compression is closer, but yeah i think thats about right.
- dr_kiszonka 3y agoI would like to better understand your edit. Can you explain why the thought divides symbols into good/bad and how it is projected to rational numbers? (What is the frequency domain of?)
- akomtu 3y agoThoughts to minds is what songs to birds: a way to communicate a relationships between observable things. Such relationships get reduced to a set of things "inside" and everything else that's outside. The boundary must be a smooth shape, that could be arbitrarily precise if the physical medium allowed it. A mind receiving a thought sees it as a finite set of resonant frequencies, even if the thought is more complex than that. Furthermore, the ultimate receiver of the thought is brain with its finite number of neurons, and even though different frequencies get mapped to different neurons, each neuron reduces amplitudes even further by activating a chemical reaction only when its input exceeds a certain threshold. Thus a continuous amplitude gets reduced to a sequnce of 0s and 1s, or a rational number. Hence the thoughts thinkable by brain are "rational". In their own realm thoughts are more like birdsongs with infinitely many harmonics and precise amplitudes.
- simonster 3y agoFor some reason, this article refers to the Chinchilla scaling laws as "data-optimal scaling laws." They are actually scaling laws that describe how to train the best model at a given computational cost, assuming that both the model size and the amount of data on which the model can be trained are constrained only by the amount of compute available. You can get an equally good model with less data if you make the model bigger, but such a model would require more compute to train than the compute-optimal model. It may also be possible to repeat the training set during training and get most of the benefits of training on more data as long as it isn't repeated too many times; this is a common thing to do in other subfields of ML but for LLMs the effect of doing so is not well-characterized.
- sinenomine 3y agoThe repeated dataset regime is better characterized than people think, see Meta's Galactica paper: https://arxiv.org/abs/2211.09085 https://arxiv.org/abs/2211.09085
- visarga 3y agoThere are three ways to train: - best score - don't care about efficiencies (GPT3, GPT4) - best score for a fixed quantity of compute at training time - good for PhD's and people who make proofs-of-concept (Chinchilla) - best score for a fixed quantity of compute at inference time - good for people who inference their models at scale (LLaMA, chatGPT turbo) The article didn't mention the LLaMA scaling laws, where we use more than 20 tokens per weight, more precisely 142 tokens per weight for LLaMA 7B. > The objective of the scaling laws from Chinchilla is to determine how to best scale the dataset and model sizes for a particular training compute budget. However, this objective disregards the inference budget, which becomes critical when serving a language model at scale. In this context, given a target level of performance, the preferred model is not the fastest to train but the fastest at inference, and although it may be cheaper to train a large model to reach a certain level of performance, a smaller one trained longer will ultimately be cheaper at inference. For instance, although Chinchilla recommends training a 10B model on 200B tokens, we find that the performance of a 7B model continues to improve even after 1T tokens. What we care about is the best model we could run on our own hardware, not how efficient was its training, that doesn't cost us users anything.
- alchemist1e9 3y agoAm I the only one who is surprised how small the storage size is? I understand it’s a lot of text but to realize it easily fits on a few cheap HDDs raided together is still incredible to me.
- alwayslikethis 3y agoText is very compressible. It can probably fit in a single consumer HDD (8TB or so) compressed with zstd.
- alchemist1e9 3y agoIt would be neat of there could be defined curated chunks of standardize text data sets for LLMs training distributed over bittorrent. Maybe there already is and I just haven’t looked hard enough.
- kangalioo 3y ago> The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together. https://pile.eleuther.ai/ https://pile.eleuther.ai/
- dougmwne 3y agoI have yet to see anyone explain or explore why text could not be trained on multiple times. And is there any reason a text passage wouldn’t be just as good run backwards as forwards? Predicting the previous word in a sentence seems just as relevant to grasping semantic meaning as the next word.
- blurbleblurble 3y agoOne of the first transformers, BERT, was and still is bidirectional.
- visarga 3y agoThey might fear unwanted memorsation/regurgitation of training content.
- ALittleLight 3y agoBut shouldn't they try it? Has someone tried it?
- thomashop 3y agoThis is already how it works. When you train a deep learning model you show the samples of the training data hundreds of times. Of course depends on the size of your training set and your budget, etc. You often train for 100s of epochs. One epooch means one pass over the training set.
- navjordj 3y agoThe original T5 models were trained with masked language modelling similar to BERT and afterwards also trained on a autoregressive task similar to GPT.
- firatsarlar 3y agoDid they read https://en.wikipedia.org/wiki/One_Hundred_Years_of_Solitude https://en.wikipedia.org/wiki/One_Hundred_Years_of_Solitude Or https://www.amazon.com/dp/0691119279/ref=as_at?tag=fivebooks001-20 https://www.amazon.com/dp/0691119279/ref=as_at?tag=fivebooks... or this guy https://tr.wikipedia.org/wiki/Ali_Nesin https://tr.wikipedia.org/wiki/Ali_Nesin Nice :)
- worldsavior 3y agoCould the text size shrink using already trained LLMs? There is probably alot of irrelevant or wrong information in this data, and the LLMs could be used to remove this information.
- astrange 3y ago> I advise government and enterprise on post-2020 AI like OpenAI’s upcoming GPT-5, and Google’s ongoing Pathways and Gemini models. OpenAI's GPT-5 that they've said they're not making yet? Does anyone actually read this guy or is this what they call puffery?
- moffkalast 3y agoWell he did say "upcoming", so it's not wrong I guess?
- SanderNL 3y agoSanderNL doesn’t know if he is just jealous, but he thinks it kind of odd to talk about one’s self on one’s own website like this: “A contributor to the fields of human intelligence and peak performance, he has held positions as chairman for Mensa International, consultant to GE and Warner Bros, and memberships with the IEEE and IET.” He doesn’t mean to degrade the guy as he doesn’t even know him. It’s just.. ah, well.
- version_five 3y agoIt's not that weird in a professional bio. For example if you read someone's bio in journal article (some ieee articles ask you for this) it will be 3rd person. Same for of you're a speaker at a conference.
- SanderNL 3y agoI know.. but at the end of every (website) article?
- kalimanzaro 3y agoTo each his antidepressant. He probably (rightly) thinks that these are more effective than an equivalent helping of SSRIs. When you cannot degrade him, pity him!
- walrus01 3y ago
- epups 3y agoI would like to understand the role of quality vs quantity in these models. Is it better to train them on slight variations of what humans consider to be high quality input or just add a boatload of low quality input instead? If we somehow found an original Shakespeare book, wouldn't it be drowned by all the trashy Amazon erotica written every year? I understand that there are reasons to expect improvement with a broader set of inputs - the model would never understand slang and other key language components by being trained only in academic papers. I wonder whether the Shakespeare example would be worked out by it somply occupyinf a novel high-dimensional space because of its uniqueness, or whether there is a signal to noise issue here.
- rationalfaith 3y ago[dead]
- pallas_athena 3y agoI like the comparison with books, but: running Whisper AI to transcribe YouTube as a whole would give us a enormous amount of data. Or would it? Are there estimations abut this?
- ofou 3y agoThis is probably what Google will do soon. We'll see!