4 ms·
Something I’ve been pondering lately: with the release of all of these content generating language models lately, what will happen to the “scrape and train” pro
by dchuk 4y ago
Something I’ve been pondering lately: with the release of all of these content generating language models lately, what will happen to the “scrape and train” process going forward? Aren’t we about to realize an infinite loop problem basically where the training sets will be full of ai generated content that isn’t actually useful for training?
Seems like we almost need pre-ai scraped datasets, almost like how it’s sometimes very useful to have metals that were created before the first atom bombs (I don’t know too much about this, just that there’s a definite land in the radiated sand around that timeframe).