2 ms·
There are quite a few papers that indicate that the important factor is the quality of the data, regardless of where it is sourced from. You still need a lot of
by throwaway4aday 3y ago
There are quite a few papers that indicate that the important factor is the quality of the data, regardless of where it is sourced from. You still need a lot of it so throwing out large chunks of your training set does work against you whether you're doing that for legal reasons or quality reasons. The solution appears to be synthetic, high quality data generated by another model but that makes it a bit of a chicken and egg problem where you need a high quality model like GPT-4 to produce high quality data reliably. I think there are methods for getting around this using less capable models by producing a lot more examples and then using another model to judge the quality of each example and only selecting the best ones. I also suspect you could get pretty far by permuting a smaller set of high quality data to produce a lot more examples that have the same meaning but differ in how they are written.