4 ms·
Some of their findings are really interesting, as well as their approach with the children stories, but I find it worrying that at no point the article mentions
by seu 3y ago
Some of their findings are really interesting, as well as their approach with the children stories, but I find it worrying that at no point the article mentions the problems inherent to the feeding AI models with the output of other AI models.
> The success of the TinyStories models also suggests a broader lesson. The standard approach to compiling training data sets involves vacuuming up text from across the internet and then filtering out the garbage. Synthetic text generated by large models could offer an alternative way to assemble high-quality data sets that wouldn’t have to be so large.
And they completely ignore the fact that those models are building their "synthetic stories" based on the knowledge from all that vacuuming text from across the internet, as if now we had already solved the problem of sourcing human data without the need for further vacuuming.
- visarga 3y agoRelated questions: What is the status of synthetic data generated from a model trained on copyright infringing content? How about the text generated from an open source model, trained on open or licensed data, but having copyrighted material in the prompt for reference. Does going through an AI model wash copyrights away? Does any hint of copyrighted data in the training corpus or prompt invalidate the right to publish the results? Only when they are similar enough to the copyrighted content? How about when the content itself is pretty common and not unique at all, like a solution to bubblesort?