4 ms·
This is wrong for at least two reasons: 1) The data scraped from the web is filtered, deduped, ranked, and cleaned in a variety of other ways before being used
by throwaway4aday 3y ago
This is wrong for at least two reasons:
1) The data scraped from the web is filtered, deduped, ranked, and cleaned in a variety of other ways before being used for training. While quantity is necessary for a model to learn the structure of language, quality is even more important once you have a model that can produce coherent output so a lot of work goes into grooming the data.
2) There have been a bunch of papers and training runs that show synthetic data created specifically for training a model is as good or better than scraped human produced data. The importance of scraped web content is quickly declining because you can now generate infinite higher quality examples using existing trained models. The only relevance it has now is for knowledge about new developments and that is a much easier stream to filter since most of the important stuff comes from official sources and you don't need as many variations since you can just generate your own using one or more examples.
- smackeyacky 3y agoHow much time is going to be wasted cleaning the data? It seems like AI can generate junk much faster than we could possibly filter it.
- throwaway4aday 3y agoThere could be an infinite amount of junk, in fact the amount of junk available online before LLM output was added to the mix was already incredibly huge. At this point it no longer matters because you can take 1 or 2 examples of some knowledge you want to train a new model on and you can then feed it to an LLM with instructions saying to create 10,000 variations on this text while still retaining the factual details, you then filter this list and throw away any output with errors or other garbage and what you're left with are a good number of high quality examples of the information you want to add to your training set. Incredibly, there is new research that shows even this might not be needed at some point because there is evidence that transformers can learn from a single example.
- s1artibartfast 3y agoyou dont need to clean all the junk created. you just need to clean enough to feed the models.
- spunker540 3y agoFortunately there are now AIs that can help with the data cleaning task.
- majewsky 3y agoBut who monitors the monitor?
- throwaway4aday 3y agoAt a certain point you just need to take a statistically representative sample of the output for review.
- smackeyacky 3y agoSeems like this is part of the feedback loop that will render AI useless though.
- throwaway4aday 3y agoNo, it's an effective technique that has been used by Microsoft Research among many others.
- ethanwillis 3y agoInfinite?
- throwaway4aday 3y agoA bit hyperbolic, I'm sure the different variations are finite but you can produce a very large number of examples with essentially the same information content.