4 ms·
That would only apply to repositories. But to train these models, you need hundreds of terabytes of diverse data from the internet. Up until now a relatively st
by mxmlnkn 4y ago
That would only apply to repositories. But to train these models, you need hundreds of terabytes of diverse data from the internet. Up until now a relatively straight-forward scraper would yield "pristine" non-AI-generated content but now you would have to filter arbitrary websites somehow. And getting the date of publication for something might be difficult or highly specific to a particular website and therefore hard to integrate into a generic crawler.