14 ms·
Will we run out of ML data? Evidence from projecting dataset size trends (2022)
- cs702 3y agoNo. AI can generate as much synthetic data as we need, on demand. Many SOTA models, in fact, are already being trained with synthetic AI-generated data. See https://en.wikipedia.org/wiki/Betteridge's_law_of_headlines https://en.wikipedia.org/wiki/Betteridge's_law_of_headlines
- version_five 3y agoLots of reasons this isn't universally true - it only works if you know enough about the data to simulate it, and your stuck within some distribution + human guesses space that's not all encompassing. The easiest counterexample is training LLMs, how are you going to synthesize useful language examples if you want more. Some version of this is true for most applications.
- mirker 3y agoYeah the issue is you can generate data, but it won’t be good data. Training over random strings won’t make you learn language, but it’s technically data.
- qeternity 3y agoYou're just sampling from an already sampled distribution. This is not the same thing. There will still be value for fine tuning, but it's no substitute.
- mirekrusin 3y agoIt's not one way consumer just like humans are not. It can direct long term evolution of reason. For starters it can be used to denoise/dedup/optimise training set to be closer to optimum (to create smaller "copies" of itself). There are instances of things that happened (history, what Paris Hilton did say on 22nd of April etc, big database of mostly irrelevant facts) and truths (math, physics, chemistry etc) where AI can enhance discoveries by helping us to see what we have not yet realised. Both seem endless tbh but personally I'm more interested in latter.
- haldujai 3y agoTo my knowledge no SOTA model has been trained on a significant proportion of synthetic data, has this changed? The best examples I know of are instruction tuning sets but that is a minute amount of data compared to the unsupervised training data.
- Buttons840 3y ago> AI can generate as much synthetic data as we need, on demand. I don't think this is right. Can I take an untrained LLM (a neural network with random parameters), and have it start generating garbage, and then train the network to produce more of the same and then have it bootstrap itself to intelligence? Of course not. What if I train it just a little bit first? What if I train it until it produces gibberish, but does occasionally string two words together that are spelled correctly. Can I have it produce petabytes of gibberish and then train on that to reach GTP4's level? You seem to argue that at some point, the AI is able to improve by training on its own output. At what point does that arrive? Because so far we've never seen an AI improve based on its own output. (As far as I know?)
- sanxiyn 3y agoAlphaZero in fact improves based on its own output, but I agree it is a special case and probably not generalizable.
- Buttons840 3y agoIt's RL though. Its output comes, in part, from interaction with an environment. It also has a well defined objective (win games). GTP doesn't have a clear objective other than "do more of this".
- blackbear_ 3y ago> Because so far we've never seen an AI improve based on its own output. Maybe it's because AI is such an overloaded term, but this is pretty commonplace for (semi-)supervised learning algorithms. Pseudo-labeling [1,2] is an example of this that has been around for decades. When done properly it does improve the performance of the original model, up to a certain limit (far from the singularity). Moreover, it is apparently possible to improve a model's performance by augmenting it's training set with synthetic examples generated by a second model [3]. Finally, boosting [4] can also be seen as iteratively leveraging the output of a model to train a slightly better model. In fact, a specific type of boosting often yields state of the art performance on tabular data. [1] https://arxiv.org/abs/2101.06329 https://arxiv.org/abs/2101.06329 [2] https://stats.stackexchange.com/questions/364584/why-does-using-pseudo-labeling-non-trivially-affect-the-results https://stats.stackexchange.com/questions/364584/why-does-us... [3] https://arxiv.org/abs/2304.08466 https://arxiv.org/abs/2304.08466 [4] https://en.m.wikipedia.org/wiki/Boosting_(machine_learning) https://en.m.wikipedia.org/wiki/Boosting_(machine_learning)
- hackerlight 3y ago> AI can generate as much synthetic data as we need, on demand. Doesn't work in majority of domains. You need to know the generating process (e.g. game rules) and build a realistic simulation environment that emulates that, in order to generate data that is useful. Both of these things are out of reach for most applications. I believe the next large step will be multi-modal, where text is contextualized by video so the LLM will be able to concretize what "sitting on a chair" actually means with a single example, without needing to see thousands of textual associations to infer the meaning from the text.
- replygirl 3y agoIf we play our cards right, AI could free people up for more valuable pursuits, and the pace of human information production would increase by orders of magnitude
- gumballindie 3y ago> free people As opposed to what? Being "captive" in jobs for paying bills?
- digdugdirk 3y agoI mean... Yes. What would you suggest as the alternative?
- gumballindie 3y agoErm not causing mass unemployment by stealing data? Also people go freely where there's pay. Seems like there aren't many opportunities and there will fewer.
- replygirl 3y agoSo hopefully we play our cards right, by extracting benefits from AI that overcompensate for the negative impacts like mass unemployment and democratization of intellectual property. If the spoils are distributed in such a way that people's standard of living is maintained or improved, people have more liesure time, which the social sciences have shown will not mean people will just stop working--they'll work less, but with higher productivity on things promising a greater benefit to family, community, and society. Forgive me if I'm misreading, but I'm having trouble with your line of reasoning. Your first reply to me scarequoting "captive" strongly implies an argument that the imperative to seek employment for survival is not a limiting factor on how people spend their time, and therefore that my suggestion that giving people more choice over how they apply their talents could be a good thing is irrelevant; but your child reply implies a concern that AI taking over some human labor will cause mass unemployment and explicitly states choice is declining. I'm advocating that, since the genie is out of the bottle, AI could be used to free people from toil, just as other labor innovations like machinery and the 40-hour work week have done. Why the dismissive snark? In the abstract, do we not want the same thing?
- brianr 3y agoThis analysis misses the impact of AI models being deployed, like is happening rapidly right now. Production applications built on AI will provide ample (infinite?) additional training data to feed back into the underlying models.
- haldujai 3y agoNot sure that synthetic or LLM-generated training data is as useful as human generated text. It seems "good enough" (for now) but synthetic makes up a very small proportion of the training set being used in current models that have been trained on it, if that proportion ends up being mostly synthetic we'll likely see whatever weird hallucinations and biases in the dominant backend (GPT4 or whatever) become amplified. It's been shown repeatedly that garbage in = garbage out for training data.
- brianr 3y agoAgree about synthetic data. My point is that AI-powered applications that are deployed in production generate more _real_ data which can be used for training. For example, self-driving cars generate tons of data about how their models perform, as a result of the cars driving around. Similarly, code-writing AI applications will generate feedback in the form of errors, logs, etc. which is can be fed back into the models as training data.
- haldujai 3y agoI wonder if the better question is not how we get more training data but: If we're running out of training data with hallucinations and performance remaining so inadequate (per OpenAI's whitepaper) is an autoregressive transformer the right architecture? Perhaps ongoing work in finetuning will take these models to the next level but ignoring the LLM hype it really does seem like things have plateaued for a while now (with expected gains from scaling).
- visarga 3y agoThere is still an order of magnitude more organic text. Ilya Sutskever recently said it was still ok. After that, we got to use reinforcement learning (agent GPTs with tools) to generate and self-validate more examples. One "simple" application would be to build a full index of facts in the whole training corpus. Just pass each document to GPT and ask it to extract the facts. Then create an inverted index, with each fact and its references. This will allow us to generate a wikipedia-like corpus of exhaustive fact research. We can say if a fact is known or not, we can tell if it is settled or controversial, and if it is a preference we can tell what is the distribution. This has got to help with factuality and generate lots of text to feed the model. Basically only costs electricity and GPU. It nicely side-steps the problem of truth by simply modelling the empirical distribution in an explicit way. At least the model won't hallucinate outside the known facts.
- haldujai 3y ago> There is still an order of magnitude more organic text. Posing this as a thought experiment, agree we still have more data to go. That we are wondering about this suggests that the current approach may be inadequate, i.e. it should not take petabytes of data for a LLM to match the performance of a high school student (for the LLM = AGI folks). > One "simple" application would be to build a full index of facts in the whole training corpus. Just pass each document to GPT and ask it to extract the facts. Agree, KG+LLM is a good next step to explore and should address some hallucination issues (see DRAGON from Leskovec and Liang groups). But we're already now talking about architectural changes as I posited. In any case, where do we get such knowledge graphs (or index of facts)? Some already exist (e.g. Wiki, UMLS) and were created by humans but are clearly inadequate in coverage. The proposition of using GPT-like models to generate these (i.e. GraphGPT) seems conceptually flawed as GPT does not itself know if a statement is factual or not which is problematic even for humans. Settled vs controversial is orders of magnitude more complex, how on earth do we do this without human annotation? You can't rely on frequency (i.e. some things were facts for 100 years but all of a sudden they're not anymore and this is not controversial by definition). The only reason LLMs work as well as they do now is because sheer volume of data (and NTP) makes the noise seem hidden and by definition an autoregressive model should be somewhat impervious to singular factoids (vs a model being grounded by the garbage dump that is CommonCrawl/the internet). > At least the model won't hallucinate outside the known facts. Not sure this is a given, even if a model acts as a natural language database of factoids it is probable that it will hallucinate links unless you're strictly grounding output in which case we've just built a colossally over-engineered IR/STS tool. > One "simple" application I think what you've posited is actually harder to build than anything that's been achieved thus far with LLMs.
- nologic01 3y agoBrute force approaches always hit some wall. ML will be no different. In the decades to come it us quite likely that algorithms will develop in directions orthogonal to current approaches. The idea that you improve performance by throwing gazillions of data into gargantuan models might be even come to be seen as laughable. Keep in mind (pun) that the only real intelligence here is us, and we are pretty good at figuring out when a tool has exhausted its utility.
- gleenn 3y agoAI had a winter of many decades because the hardware wasn't there and there were better alternatives, especially for neural nets. Now ChatGPT etc comes out, with unbelievable results, decades in the making. And a couple months we're already writing it off because of the next limitation? Maybe let's give it more than a month or two to figure out if we even need all that data. I heard they're already talking about trying to significantly reduce the model hyper parameters size even though a large model size increase apparently the reason GPT 4 was so much better than 3. Give it a minute IMHO before making generalizations like this so soon
- mjburgess 3y agoWell I imagine the commenter actually understands the domain, the techniques, and is making an informed opinion. It is possible to form opinions by knowing the domain, rather than drawing an exponential curve of newspaper headlines which trails off "..."
- airgapstopgap 3y agoWe won't hit the wall. Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an OOM more text tokens. But there are other sources of data, and slightly different ways to utilize it. Multimodality, in very large training runs, will almost inevitably increase sample efficiency (for obvious reasons of context richness), synthetic data is already very effective [1], and there are and will be discovered other ways to do more in the condition of diminishing raw text resources. But a thorough abandonment of the scaling strategy is very unlikely. Sutton's Bitter Lesson [2] points at a very powerful rule of thumb: we shouldn't turn AI engineering into a contest of smartness, we should allow complex smartness to emerge from generic low-level algorithms. What will be seen as laughable in decades to come is not the scaling strategy, but the Godlike conceit of people who thought they can devise generally applicable rules of reasoning from first principles. 1: https://arxiv.org/abs/2304.08466 https://arxiv.org/abs/2304.08466 2: http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- flyval 3y agoThis is dumb. An individual human takes in more data than modern LLMs do. https://open.substack.com/pub/echoesofid/p/why-llms-struggle-with-basics-too?utm_source=direct&utm_campaign=post&utm_medium=web https://open.substack.com/pub/echoesofid/p/why-llms-struggle...
- majikaja 3y ago>Just the vision data of a baby’s first year easily adds up to petabytes What encoding is this??
- lostmsu 3y agoUncompressed 2x 8k by 8k 24bpp 24FPS video. Comes at about 500GB per hour.
- majikaja 3y agoMaybe if we store text data as sequences of 10k x 10k PNGs (one for each letter) and add an image recognition layer it would improve LLM perf
- amscanne 3y agoIt’s always interesting to think about the exact technological analog for our biological sensors, but I believe that our vision would be way less than that in terms of raw data. We have a super high-res area at the center of vision (the fovea), but the rest is extremely low resolution (but with high movement and light sensitivity). I think we could reasonably say that if an optical nerve has 1mm neurons on average, and they can fire at 250Hz at the most, that’s 250mbps or ~31mb/s per eye of uncompressed data as an upper bound.
- eastbound 3y ago31mb/s = 14GB/hr (bits to bytes). 81TB per year, assuming 16 hours awake per day. Fits snuggly on a large SSD ;)
- 3y ago
- laserbeam 3y agoI love how "running out of data" implies that AI companies have access to all the text we ever wrote on all platforms out there. I mean it's probably true...
- StrangeATractor 3y agoOn this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.
- albert_e 3y ago100% -- I also see this as a big and emerging problem that future researchers and practitioners will have to deal with. Posted some thoughts previously here -- https://news.ycombinator.com/item?id=32577822 https://news.ycombinator.com/item?id=32577822 https://news.ycombinator.com/item?id=33869402 https://news.ycombinator.com/item?id=33869402
- Qem 3y ago> as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction. If we build a system where we feed the exhaust of an AI to another one at each step, should we call it the AI Centipede, like in the movies? https://m.imdb.com/list/ls064583741/ https://m.imdb.com/list/ls064583741/
- furyofantares 3y agoI see this take a lot and I think it's quite wrong, not fully, but at least missing a couple big points. I think they already don't blindly feed it just all the garbage raw data they can find, but prefer high quality, well-prepared sources. And aside from spam, we're not just blindly posting AI content either. We're putting in meaningful prompts, rejecting answers we don't like, and editing answers we do.
- te_chris 3y agoUm, really? You think your average 'growth hacker' who is using ChatGPT to exponentially increase the amount of SEO junk they can churn out is checking each answer before they press publish? Purity, accuracy and relevance of data collected from the internet is going to a very hard problem.
- bhouston 3y agoWe just are not thinking wide enough: * Train on all of television history, and streaming content. * Train on YouTube. * I suspect at some point we'll have a recording of most of people's lives, e.g. live-streaming: https://en.wikipedia.org/wiki/Lifestreaming#Lifecasting https://en.wikipedia.org/wiki/Lifestreaming#Lifecasting
- mountainriver 3y agoExactly, put bots into the world with cameras and you have infinite training. Humans also need a ton of data to train on and have way more parameters than the biggest ML model today
- eru 3y agoYou can also gather arbitrarily more video data by just turning on some webcams and pointing them at the world. In addition you can also feed your system from video games.
- joseph_grobbles 3y ago[dead]
- bobsmooth 3y agoThere's gotta be entire libraries that haven't been digitized that can be mined for data.
- MagicMoonlight 3y agoOnly if you rely on dumb learning where it’s learning pure pattern matching rather than interacting and reinforcement learning based on the responses.
- pixelmonkey 3y agoThe last gen of popular LLMs focuses on publicly accessible web text. But we have lots of other sources of "latent" or "hidden" text. For example, OpenAI's Whisper model can turn audio to text reliably. If you point Whisper at the world's podcasts, that's a whole new source of conversational text. If you point Whisper at YouTube, that's a whole new source of all sorts of text. And then there are all sorts of private sources of text, like UpToDate for doctors, LexisNexis for lawyers, and so forth. I suspect "running out" isn't a within-a-decade concern, especially since text or text-equivalent data grows exponentially in the present internet environment. I think the bigger challenge will be distinguishing human-generated from AI-generated data after 2023.
- Salgat 3y agoBut how much more data is required to make a big difference? Is doubling the dataset considered a dramatic improvement? Or is increasing the dataset by 10x needed?
- ospray 3y agoAlso quality is likely important will the models get better if we train them on YouTube comments.
- chii 3y ago> If you point Whisper at YouTube, that's a whole new source of all sorts of text. a lot of YT videos already has autogenerated english subtitles, which is actually available as a vtt download, so don't even need to use Whisper on a video to obtain it!
- HybridCurve 3y agoThis take is a bit silly in that they are implying the problem training models will be that we will run out of data. It's more likely that the problem is that the current models require too much data to reach convergence. We've been trying to speed run neural networks science for the past decade but we still don't fully understand how they work. It's like being a bad programmer who doesn't understand algorithms so you compensate by spending money on hardware to make your programs run faster. At some point we will reach a limit where you can't buy your way out of the problem with more data or money and we'll all be forced to return to studying the foundations of the science rather than just trying to scale the existing models up. I am certain when we get to that point everyone will realize we've been trying to feed these models too much data. It makes more sense that our current architectures are just not effective at assimilating the data they have.