5 ms·
There's not enough human workers to validate that at scale. If you want ML to do it... well that's a bit of a catch-22. How would an ML algorithm know if data
by devmor 2y ago
There's not enough human workers to validate that at scale.
If you want ML to do it... well that's a bit of a catch-22. How would an ML algorithm know if data is good enough to be trained on unless it has already been trained on that data?
- HarHarVeryFunny 2y agoThe way humans do it is via curiosity/boredom/surprise. If we can't predict some thing well (i.e. we're "surprised" by it) then that both acts as learning trigger to predict better next time, as well as retains our interest to explore it. Eventually AGI's will need to learn by experimentation as we do, but in the meantime ability to well predict a potential training sample could be used to decide whether to add it to training set or not. At the moment it seems the main emphasis is on a combination of multi-modality (esp. video) and synthetic data where one generation of LLM generates tailored training samples for the next generation. I guess this synthetic data allows a more selective acquisition of knowledge than just adding surprising texts found in the wild.
- IncreasePosts 2y agoYou could use data where you know it was not AI generated, like the library of Congress catalog from prior to 2015. Or highly cited research papers, things of that nature.
- drdeca 2y agoGet a collection of data which is small enough to have humans annotate as being either low quality or high quality. Train a model to predict this annotation. Then on a larger disjoint collection of data, use this model to estimate whether the data points would be considered low quality or high quality, and use this to filter it. This seems doable, and, I think something like it is already done?