6 ms·
Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data
- thanhtan 2y ago[flagged]
- ofrzeta 2y ago"If trends continue, language models will fully utilize this stock between 2026 and 2032" - that will require data centers with their own nuclear reactors (or other power plants) as hinted at by Marc Zuckerberg?
- trott 2y agoIf you take Llama-3-400B, and 30x its data (hitting the data ceiling, AFAICT), 30x its size to match, and the hardware improves by, say, 3x, then you'll use up about a year's worth of energy from a typical nuclear power plant.
- mathsmath 2y agoI don’t know much about LLMs, but is it possible to throttle their training? Solar has gotten pretty cheap, and I’m just wondering if you can throttle up and down based on how much output the panels are producing.
- moi2388 2y agoOf course it is, but the trade-off is time.
- spiralk 2y agoIf its for training a new foundation model it is not that bad. It's still only a fraction of the energy compared to many human industries. I did rough math some time ago and found that that training llama-3-70B used the equivalent energy to 1/30 of a full loaded container ship going from China to the US. Even scaled up 100x and trained 10x longer, its seems like the energy consumption is relatively small compared to other industries. The fact that people are considering nuclear power for AI training is an advantage not a downside, imo. It should have a much lower CO2 footprint.
- adrianN 2y agoYou always have to compare the cost to the value it generates. A year of power from a nuclear plant might be used in more productive ways.
- ofrzeta 2y agoYes, and consider that with the current hype around AI and enough (venture) capital there will be several corporations competing for the best AI and suddenly we are at several "nuclear plants" or equivalent energy "consumption".
- spiralk 2y agoI don't see increasing demand for nuclear power as a disadvantage. We have nuclear material that can last humanity 1000s of years at least. The CO2 footprint is an issue but nuclear is much better than others. Personally, I think it's better we utilize more energy and discover new breakthroughs while society is relatively stable and functioning, because there's no guarantee that it will last. Population collapse seems imminent in more educated societies, even China and India are trending this way now. Without some level of AI assistance, humanity would likely lose a great deal of productive output. Also, if this path to AGI does not work out, its not as though the nuclear reactors will be wasted. People will find something else to do with the energy.
- spiralk 2y agoSure I agree, but if we compared value it generates per unit energy it would still probably be better than many non-essential industries: the entertainment industry, fashion industry, alcohol, etc. Even in the current state LLMs can provide more useful practical value compared to industries with higher energy and CO2 footprints.
- monero-xmr 2y agoIf someone is willing to pay, who cares? Energy has a price. Focus on regulating how energy is generated, and when prices climb the market will solve the problem. If instead you focus on using the government to outlaw demand, only failure will follow. I mean, didn't the government outlawing the demand for illegal drugs fail miserably? I believe drugs are cheaper, more potent, and more available than ever. Similarly, if there is demand for compute, then compute will occur. There is always a clearing price commiserate with the risk.
- amanaplanacanal 2y agoA carbon tax would be the most free market way to do it: tax fossil carbon as it comes out of the ground. The market can handle the rest. Can’t seem to make that happen politically, though.
- DrNosferatu 2y agoCareful with carbon tax plans, they can be regressive: https://blogs.worldbank.org/en/energy/what-carbon-tax-can-do-and-why-it-cannot-do-it-all https://blogs.worldbank.org/en/energy/what-carbon-tax-can-do...
- bamboozled 2y agoRemember tackling climate change, Remember all the Silicon Valleys pushing for us to tackle climate change?
- surfingdino 2y agoYeah, where's that app that was supposed to fix it?
- throwaway48476 2y agoWere not even close to running out of human generated data. The reason it seems this way is because it's so hard to find old data. There are tons of whole magazine scans on some obscure website that's not even indexed. Most of this is the fault of Google who has been an atrocious steward of search. Why is it that I still can't do full text search of the internet archive dataset? Forever copyright of commercially de minimus works also plays a large role. There's a monumental amount of quality data out there that's not indexed, not searchable, and abandoned but unused. We just need to value it enough to use it.
- stubish 2y agoSo much old, out of date, factually incorrect, racist, sexist and even illegal information. I think it is already clear that training systems on everything is not the way forward, and about as reliable as the set of 80s Encyclopedias my mother refuses to throw out. The current tech needs to be trained on good data to produce good results, as it can't reason and gauge reliability or even pick up when its output is self contradictory.
- nobutterbetter 2y ago[dead]
- glimshe 2y agoThe geniuses and stewards of our civilization of just a couple of decades ago were trained on this very data. We don't yet know what outcome we'll get by handing out the world to the people trained on "new, up to date, factually correct, egalitarian and legal" data.
- stubish 2y agoWe hope they used their reason to maintain their knowledge over the years, or at least updated their poor fashion choices. Or maybe not given so much effort is made to enforce moral opinions from biblical times.
- 2y ago
- _boffin_ 2y agoThe amount of data that all the different government agencies has tucked away in their different file cabinets has to be magnitudes more than what's on the public internet. The amount of data in the military... i couldn't even fathom. One data source i've been thinking about that i don't know if they've hit yet is all the different agencies local and state agencies and their private and public meetings, ordinances, discourse, etc...
- freilanzer 2y ago> The amount of data that all the different government agencies has tucked away in their different file cabinets has to be magnitudes more than what's on the public internet. The amount of data in the military... i couldn't even fathom. Definitely not when it comes to text. The internet is the largest resource. I'd like to see all books in the Vatican digitalised, if they're not already - probably not though.
- dgoodell 2y agoAs someone who works for the nasa, I’m not so sure. You’d be surprised how much stuff gets randomly thrown away to save space. And it’s going to get worse I now that paper files are disappearing. I wanted some info and data from a test we did 9 years ago. It was a pretty big deal, lots of people involved, many millions of dollars, multiple nasa centers contributing. Every single person on the test randomly kept their own files for the portion of the test they were responsible for. And the only copy of the raw test data was deleted by one of them to save some space when upgrading. There is no record anywhere of what equipment was used for the test. One of my coworkers has 4 TB external HDD that he keeps everything he has ever worked on. It’s not backed up anywhere else. It just failed and he thought he lost everything, luckily I was able to recover most of it. Wtf.
- makapuf 2y agoFunny that it does not need that much data to train your average 20th century human genius. I'd say that if we are dreaming of the future of ai, learning and reasoning seems the greatest issue, not data. That said, the article title is about LLMs, so that's what will need changing I guess.
- jstanley 2y agoHumans aren't just text interfaces though. The majority of your input is not textual but is sights, sounds, feelings, etc., that LLMs don't (yet?) have access to. Humans receive an enormous amount of training data in forms not currently available to LLMs. If you locked baby Einstein in a room with the collected works of humanity and left him there for a lifetime, I doubt he'd have even learnt to read on his own.
- nope1000 2y agoWe also inherit a lot of network topology already
- devoutsalsa 2y agoI remember listening to an AI researched in some interview over 20 years ago. He said that in his quest to create an AI, he realized at some point he could just have kids instead.
- mrweasel 2y agoWe also have a funny way of applying solutions and lessons learn in one field to problems in completely unrelated areas. Given the statistical nature of LLMs I'm not convinced that they are able look across fields in the same way as a human brain, they lack creativity. The greatest advantage you can have in life is a creative mind and I don't believe that is something that can be taught. It can be stomped out of you as a child, but it's not learnable.
- lassoiat 2y agoI have come to the point that it is not really fair to the LLM to statistically train it on human output and expect it to come up with something more than the average. There will be much value in automating the tedious and the routine. Of course, that doesn't make for a great science fiction story. We first have to placate all these science fiction fantasies and in the process we will automate the tedious and the routine as a side effect of trying to figure out how many AGI can dance on the head of a pin. Then human creativity will just be worth all the more.
- aaron695 2y agoThey haven't even translated non-English material and mixed it all in yet (that I know of) This is big because it would hold novel data the West doesn't access. What is the 'mood' of the average Chinese farmer on Taiwan. Otherwise it's hard to see how adding more text of the same thing is going to create a revolution. Video will be something new. But if like "Her" it watches every Twitch stream simultaneously for a month, and is talking to a billion people for a month and still doesn't get it what else is going to happen?
- surfingdino 2y ago> They haven't even translated non-English material and mixed it all in yet (that I know of) The current performance of LLMs on non-English languages is disappointing. Feeding it more non-English material is not going guaranteed to help. > This is big because it would hold novel data the West doesn't access. What is the 'mood' of the average Chinese farmer on Taiwan. The average Chinese farmer does not produce textual output of that kind. It is generally not advisable to put your thoughts in writing in oppressive regimes. It could be a life-ending mistake. > Otherwise it's hard to see how adding more text of the same thing is going to create a revolution. The LLM gang are like the people who think they can get slimmer by eating more. > Video will be something new. But if like "Her" it watches every Twitch stream simultaneously for a month, and is talking to a billion people for a month and still doesn't get it what else is going to happen? Since when does Twitch carry broadcasts that have any value to humanity? Is it used to hold scientific discussions? Or for shooting shit and pushing paid products and services?
- bambax 2y agoThe paradox is that the amount of data available for LLM training is going down, not up, because earlier models made ample use of copyrighted works that later models won't have access to.
- LoganDark 2y agoNot only that, but a dataset that includes LLM-generated content has been known to reduce model quality. I remember there being a paper on it but I can't seem to find it now. Essentially, the internet now being chock full of LLM garbage means that any model you train on it is going to end up quite a bit worse than it could have been, simply because of the dataset being "poisoned" by preexisting LLMs. I bet OpenAI's only real advantage is having a dataset that was gathered before LLM use was widespread.
- fifteen1506 2y agoI thought having the ex-NSA chief on board would mean a military AI would have lots of info to be fed on, in near real-time.
- DrNosferatu 2y agoAlso, important to keep in mind the inevitable contamination with AI-generated content in datasets from now on.
- resiros 2y agoLots of assumption here. First, that we will only be training on text data, if we take into considerations all the videos and audios shared I am quite sure we would have one or two orders of magnitude more of data. Second, that it even matter, there has been some early research showing that training on the right data improves prediction more than training on more data (which intuitively makes sense, training on papers and book is much more useful than training on youtube comments). Additionally, lots of the improvement in quality are because of RLHF, which is basically manual human labeling. And last, my guess is that improvements in architecture are what will unlock the next level of performance, not just scaling.
- trott 2y ago> Lots of assumption here. First, that we will only be training on text data, if we take into considerations all the videos and audios shared I am quite sure we would have one or two orders of magnitude more of data. 1GB of text is way more useful for generating text than 1GB of video is. > training on the right data improves prediction more than training on more data Books are more useful than Facebook rants. But this is an argument for data scarcity rather than for data abundance.
- nojvek 2y agoArgument of running out of data is kind of stupid. We have billions of cameras, microphones and IMU/GPS sensors. In-fact one in almost every pocket and desk. Survival requires intelligence being energy and resource efficient. Those who build the most powerful and useful models that run locally on edge and are data efficient have a higher chance of winning. Whoever provides the cheapest fastest most useful models will keep on winning.