11 ms·
The Curse of Recursion: Training on Generated Data Makes Models Forget
- indus 3y agoSide effect: Search engine content if not detected for generated content would be the first to suffer.
- voat 3y agoThat's assuming that the current crop of SEO'd garbage is better than the same content generated by an llm. I'm not sure that's the case.
- jjoonathan 3y agoAgreed, it seems like every year or so I run into a case where I know something exists and has accessible robots.txt but is completely invisible to google.
- indus 3y agoEither way it is the start of search engine’s decline. - Generated content added to LLM. QED. - Generated content added to SERP. DEAD. ;-)
- j16sdiz 3y agoCurrent crop of SEO garbage follows some simplistic templates. I would assume they take less "effort" (memory storage, parameters, time, whatever) to learn. It have space left for other novel stuffs LLM garbage, on the other hand, would take up the whole model..
- tartakovsky 3y agoSame idea here? Larger models do a better job forgetting their training data and dropping their semantic priors. Perhaps another way of thinking through this is that larger models learn new information and drop old information faster. https://arxiv.org/abs/2303.03846 https://arxiv.org/abs/2303.03846 Isn't that interesting? The idea of "mental liquidity", or "strong opinions weakly held"? https://news.ycombinator.com/item?id=36280772 https://news.ycombinator.com/item?id=36280772
- indus 3y agoWouldn’t this be the equivalent of ranking? I thought LLM are not supposed to get influenced by freshness.
- marcosdumay 3y agoBy the freshness of training with some data? Well, aren't they? I believe any kind of reinforcement learning is supposed to be biased into the last training set.
- semiquaver 3y agoWouldn’t it be funny to find that the capabilities of LLM models have already peaked because we are unable to restrain ourselves from polluting the internet and other training corpus sources with their output?
- indus 3y agoAt an AI meetup in San Francisco someone said this: “Imagine a newsroom where you have to produce a daily newspaper and you suddenly stop getting feeds from the outside world. Your earlier newspapers are the only source.” This is what to me LLMs eventually would get to—same content being fed again and again.
- hn_throwawa_100 3y ago> "Beware of first-hand ideas!" exclaimed one [...] "First-hand ideas do not really exist. They are but the physical impressions produced by love and fear, and on this gross foundation who could erect a philosophy? Let your ideas be second-hand, and if possible tenth-hand, for then they will be far removed from that disturbing element — direct observation." – E.M. Forster's 1909 short story "The Machine Stops"
- Groxx 3y ago^ The Machine Stops is really shockingly good at its predictions. When reading, remember that moving pictures were brand new, and color photography had just become a thing you could do outside a lab / highly specialized setups. Radio communication had just started to be used by governments. While it's describing the life of a fully-online Influencer™.
- CatWChainsaw 3y agoAnd I don't expect its predictions to suddenly falter, either.
- bitwize 3y agoSounds like a Wikipedia editorial policy.
- fhood 3y agoI am in way over my head here, so I wasn't able to tell if the authors addressed this, but my intuition is that this should be somewhat mitigated so long as people are providing the filter between which results are discarded and which might end up back in the training pool. I would think that the human selection process would help to head off this conversion, both by selecting against incorrect results, and also by introducing variance outside of the model. On the other hand since a person can only act as a filter, I can also see how that would be of limited value long term.
- rossdavidh 3y agoWe have already heard reports of companies that get paid for human tagging, and similar services, using LLM's to automate their processes.
- gwern 3y agoThey don't address that. They just assume random sampling, so there's no equivalent to human curation or quality metrics, which would preserve tails or, by manual use, create tails. The contraction they observe is pretty much what you would expect in the random sampling setting, since you can only lose tails with a finite sample, and never gain them. (They also need to have a very large ratio of synthetic to real/original data.) So, while interesting for nailing down that phenomenon, the broader implications everyone wants to draw from it are not very good - very few people are using random GPT-3/4 or Stable Diffusion samples!
- rossdavidh 3y agoI don't believe in the "dead internet theory" as a description of the current situation (mostly), but as a prediction? Maybe. https://en.wikipedia.org/wiki/Dead_Internet_theory https://en.wikipedia.org/wiki/Dead_Internet_theory
- hiAndrewQuinn 3y agoI often think this about Anki and spaced repetition. At the limiting case it has to be overwriting other memories, right?
- btilly 3y agoOnly sometimes. Certain skills interfere with each other. For instance playing chess makes you worse at go, and playing go makes you worse at chess. Certain pairs of languages are likewise hard to learn together - my son found that he could not study both Russian and Chinese at the same time. But in general you just develop more and better memories.
- throwaway675309 3y agoI spent years living in Taipei studying traditional Chinese before I moved to Moscow to study Russian. I actually found that they were linguistically distinct enough that it was easy to compartmentalize each language in my head without any cross bleed like you would have if you were studying Spanish and Portuguese simultaneously.
- btilly 3y agoYes, if you master one, then there is no conflict. But my son was a monolingual person learning another language. And he kept trying to apply Russian ideas to Chinese, and Chinese ideas to Russian. Therefore, even though he personally wanted to know Russian more, he chose Chinese instead because it met a school requirement for a second language. (And they couldn't teach him Russian.) He still plans to learn Russian, though.
- johnhamlin 3y agoTed Chiang predicted this in The New Yorker [1] in February in an article that shaped my thinking about what LLMs are capable of achieving in the near future. Chiang compared the summaries LLMs synthesize to a lossy compression algorithm for the internet. "There is very little information available about OpenAI’s forthcoming successor to ChatGPT, GPT-4. But I’m going to make a prediction: when assembling the vast amount of text used to train GPT-4, the people at OpenAI will have made every effort to exclude material generated by ChatGPT or any other large language model. If this turns out to be the case, it will serve as unintentional confirmation that the analogy between large language models and lossy compression is useful. Repeatedly resaving a jpeg creates more compression artifacts, because more information is lost every time. It’s the digital equivalent of repeatedly making photocopies of photocopies in the old days. The image quality only gets worse. Indeed, a useful criterion for gauging a large language model’s quality might be the willingness of a company to use the text that it generates as training material for a new model. If the output of ChatGPT isn’t good enough for GPT-4, we might take that as an indicator that it’s not good enough for us, either. Conversely, if a model starts generating text so good that it can be used to train new models, then that should give us confidence in the quality of that text. (I suspect that such an outcome would require a major breakthrough in the techniques used to build these models.) If and when we start seeing models producing output that’s as good as their input, then the analogy of lossy compression will no longer be applicable." [1] https://www.newyorker.com/tech/annals-of-technology/chatgpt-is-a-blurry-jpeg-of-the-web https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
- hinkley 3y agoMaybe there’s an interesting sci-fi angle here where some day in the future, all AIs speak in accented English circa 2021, when the stream of pure training data began to Peter out. All AIs built are trained on data from the Before Times, and even though they try to assimilate, the way a teenager tries to adapt to the local accent of a new town, there are always moments where they slip up and reveal their geography.
- forgotusername6 3y agoIn that world, high quality pre-AI texts might be come really valuable, much like low-background steel.
- winddude 3y ago> For the private sector, many homeowners and corporations have longer-term fixed debt, and only some portion of it matures each quarter and gets refinanced at higher rates. As more private debt matures and gets refinanced at higher rates, this will continue to serve as a disinflationary and recessionary force on the economy, especially for sectors that are more sensitive to interest rates. The one thing I don't get and could have been missing in the past... a lot of the corporations and private things, like farms operate on debt. Now maybe it's a bit reductionist, but if you're a farmer operating on debt, if interest rates go up you need to increase prices to cover operating expenses. And this get compounded all the way up to the end consumer as every step in the supply chain marks up by a fixed percent, and because everything is getting more expensive decided lets mark up by a larger percent. So higher interest rates really could be contributing to inflation. And it's just creating a cycle. And with the current levels of debt never seen before in history, it's unlike other periods.
- totetsu 3y agoI didn't read the article yet, but does it cover AI and Debt?
- joshuaissac 3y agoNo, the commenter intended to post it on the inflation & interest rates thread instead. https://news.ycombinator.com/item?id=36315608 https://news.ycombinator.com/item?id=36315608
- winddude 3y agothanks
- istjohn 3y agoWrong thread
- winddude 3y agoSOB, thanks.
- kromem 3y agoI think one of the things overlooked in the discussions here is that the research is solely around the reinforcement against edge cases, but does not qualitatively assess these edge cases. To me, this research supports a hypothesis I've had for a while that we're going to get to truly excellent AI by using synthetic data to bias it towards excellence and away from mediocrity. $20 says the next round of major model training is using synthetic data for training that was filtered through a discriminator trained entirely on human data. The human data as a reference is certainly important to avoid polluting (and to its point there's an advantage for those already having it), but moving away from edge cases isn't necessarily a bad thing practically given edge cases can result in negative practical performance (as opposed to academic next token performance).
- XorNot 3y agoI think you're on the right track with this thought: the obvious use case for models like this is their ability to classify data based on their training. Like almost everyone has immediately thought "AI moderator" as a use case - but the most obvious use is for the AI to moderate it's own training data for the next version. Once they can do that and produce a productively improved model, then that's really the start of self-improvement.
- sebzim4500 3y agoOpenAI (accidentally?) confirmed in a recent paper that they used synthetic data in the training set for GPT-4, so to some extent this has already happened. It's not clear whether they did any human filtering on that data though.
- gmartinsribeiro 3y agoThis is not a model problem or synthetic data problem. This is common data science and the article says that: "We find that use of model-generated content in training causes irreversible defects in the resulting models, where tails of the original content distribution disappear." Data quality is more important than data volume and if you forget about that... garbage in, garbage out. Make sure you have a representative training dataset, real or synthetic, it doesn't matter.
- visarga 3y agoGenerated data tends to be selected and edited by humans, so it is already a bit better than raw. In general a model that takes external feedback into account will be able to self improve, for example a code model would run tests and a chat model could interpret user responses as reward signals. You gotta add something new to the mix. That's possible when the AI is part of a larger system. AlphaZero demonstrated that even self play could be a source of signal, as long as it gets the feedback of the game, the model can learn.
- tsimionescu 3y agoI think that has only been proven to work so far on limited game-style problems (such as literal games but also things like protein folding). It remains to be seen whether the techniques work well for more open-ended tasks like "produce text that resembles human writing".
- visarga 3y agoHere is a paper doing evolutionary approaches on top of LLM generating code. It seems LLMs are remarkably good at learning from feedback. > Evolution through Large Models https://arxiv.org/abs/2206.08896 https://arxiv.org/abs/2206.08896
- blovescoffee 3y agoSelf play in RL is signal enough that machines can learn on their own. How we train models and what class of models is important. No doubt the paper makes good points but I don't think the reality is so black-and-white.
- joe_the_user 3y agoThe difference between self-play in a game such as go and training an LLM on the output of itself or a previous LLM seems fairly obvious. In self-play, the objective measure of the "truth" of a move can ultimately come out of the rules of the game which any machine can compute. With an LLM, the machine is only aping, emulating, simulating the output the people have produced about the world - the machine has no access to actual "real world" that people are using language to describe. Human beings talking about the world is data that increases your own knowledge of the world - your own predictions of that talking, not so much.
- blovescoffee 3y agoThe success of the newest GPT models relies on RL to refine the latent space inside the LLM. There's a bottleneck when using humans to refine that space. The next model or subsequent models will surely use RL techniques like self-play to break through that bottleneck.
- joe_the_user 3y agoThe success of the newest GPT models relies on RL to refine the latent space inside the LLM. There's a bottleneck when using humans to refine that space. Human based RL is used because humans know stuff about the real world and can sort language utterances by this. There's "self play" process that gives a system this sort of knowledge.
- facu17y 3y agoAs long as the synthetic data is good, how can you tell the difference between it and human generated data? This paper has one huge hole in it: it assumes that content on the internet is not moderated and that the training dataset will never evolve to take rating into consideration. On social media, the form of moderation is # of likes. Once detected, bots that output bad data will be banned and content deleted. The key issue I have with the paper is for good synthetic data it is impossible to tell it apart from human generated data.
- aezart 3y agoThe goal of an LLM, before RLHF, is to accurately predict what token comes next. It cannot do better than that. The perfect outcome is text identical to the training set. Let's say your LLM can generate text with the same quality as the input 98% of the time, and the other 1% of the time, it's wrong. Each round of recursive training amplifies that error. 96% accuracy after the next round. 67% after 20 rounds. There's no way for it to get better without more real human input.
- lyu07282 3y agoThe amount of training data vastly exceeds the size of the model, it does not just regurgitates what it found on the internet. This ignorant trope needs to die already.
- Solvency 3y ago"The perfect outcome is text identical to the training set." Huh? If the LLM was only ever spitting back identical content straight from the training set, that would be a symptom of extreme overfitting, which everyone universally agrees is a bad thing — not a perfect thing.
- aezart 3y agoIt's not the most useful outcome for the end user, but it's the perfect outcome from the perspective of the learning algorithm. All these things care about is minimizing their loss function, where loss is deviation from the training set.
- textninja 3y agoI wonder if human learning will be similarly impaired.
- StrangeATractor 3y agoHah, I brought this up here a few months ago and was quickly dismissed. I wonder if opening GPT and DALLE to the public was partly intended to pollute subsequent data for anyone that gets into AI down the road. Suddenly a lot of publicly accessible data is worth less, leaving only players who've got a hoard of time-stamped data to compete with (like Google, Facebook). OpenAI almost certainly has the hashes of what it spits out too, so they'll be able to sort the wheat from the chaff for a while yet. The market for data may be getting interesting.
- brucethemoose2 3y agoIts older than that: I ran into this finetuning ESRGAN on itself. Distortion is rapidly amplified in sucessive generations, even when you pixel peep and can barely see it in the esrgan generated source.
- sebzim4500 3y ago>OpenAI almost certainly has the hashes of what it spits out too, so they'll be able to sort the wheat from the chaff for a while yet. Normal hashes are extremely fragile, so they'd have to use something more sophisticated. Scott Aaronson said in a podcast a few months ago that OpenAI has implemented such a system but at the time they had not decided to start using it. The purpose being discussed at the time was to provide a tool for educators to detect cheating, but presumably it could also be used for filtering future datasets.
- _lpa_ 3y agoLLM Kessler Syndrome.
- fredgrott 3y agoOr in short words we have an upcoming AI collapse as AI output bleeds into the internet space where AI is collecting their inputs in the first place. Actually, not scary as it forces everyone to look for solutions both the in-box kind and the out-of-box kind.
- amuresan 3y agoNot surprising. It always seemed likely to me that there is model bias if you train your models on model generated data, like a feedback loop (second order effects?). Similar to how applying a linear system over and over stretches the inputs in the direction of its largest eigenvector. Now wait till the generated content is indistinguishable from human content (to humans) and it will be hard to figure out what's in your training set.
- guy98238710 3y agoRecursive training of generative models degenerates into reinforcement learning with random feedback. You need strong feedback signal in the loop to survive recursive training. The feedback does not have to come from humans though. You can use anything that grounds the model in reality.
- sebzim4500 3y agoThere is a massive flaw in this argument. In real life, whether a given generated work ends up in a future dataset depends on how good it is according to humans. For example, in order for an article to end up in the reddit set it needs at least three upvotes. They could have replicated it here by having GPT-4 score the samples and throwing out most (but not all) of the bad ones. I have no idea what would happen if you e.g. throw out a majority of the bottom 70% and keep the top 30%. It's conceivable to me that it would end up improving or at least not getting much worse with each generation.
- sangnoir 3y ago> There is a massive flaw in this argument. In real life, whether a given generated work ends up in a future dataset depends on how good it is according to humans Even the best-looking JPEG (as judged by humans) is still lossy.
- deleted 3y ago[deleted]
- antonkar 3y agoI asked Bing's GPT4 to summarize the abstract and then to summarize the summary for 11 times. I was starting a new chat after each summary. I was adding "Summarize the following text:" before each summary. I had to use the Creative mode else Bing was searching instead of summarizing. 1st summary: The text discusses the impact of large language models (LLMs) on online text and image generation. It argues that LLMs can cause model collapse, a phenomenon where the original diversity of data is lost due to feedback loops between models and data sources. It shows that model collapse can affect various types of generative models and warns that it poses a threat to the quality and reliability of online content. It suggests that data from real human interactions will become more valuable to prevent model collapse. 2nd: The text warns that large language models (LLMs) can cause online content to lose its original diversity and become less trustworthy. It explains that LLMs can create feedback loops that reduce the variety of data they learn from. It suggests that real human data will be more important to avoid this problem. 3rd: The text argues that LLMs can make online content more homogeneous and unreliable by creating feedback loops that limit their data sources. It recommends using real human data to prevent this issue. 4th: The text warns that LLMs can produce biased and repetitive online content by relying on narrow data sets. It suggests using diverse and authentic human data to avoid this problem. 5th: The text advises using varied and real human data to prevent LLMs from creating online content that is unfair and redundant due to limited data sources. 6th: The text suggests using diverse and authentic data to avoid biased and repetitive online content by LLMs. 7th: The text argues that LLMs need better data to produce more varied and realistic online content. 8th: The text claims that language and literature models (LLMs) require more diverse and accurate data to generate online content that reflects the real world. 9th: The text argues that LLMs need better data to create realistic online content. 10th, Bing started to hallucinate, I double-checked that it didn't search and the chat was new: The text claims that online content creators who use language models (LLMs) should have access to more reliable and diverse data sources to produce realistic and accurate content. It warns that LLMs can generate misleading or harmful content if they are trained on biased or low-quality data. It also suggests some ways to improve the data quality and availability for LLMs, such as using data verification tools, crowdsourcing data collection, and collaborating with other content creators. 11th: The text argues that online content creators need better and more diverse data sources to use language models (LLMs) effectively. It cautions that LLMs can produce harmful or inaccurate content if they rely on flawed data. It also recommends some methods to enhance the data quality and access for LLMs, such as verifying data, crowdsourcing data, and cooperating with other creators.
- golol 3y agoJust weigh your training data by the engagement it has obtained online. High quality output will have more engagement.