2 ms·
Labs spend billions hiring experts to generate new data, and better models can better filter existing training data and generate new synthetic data. There’s no
by sothatsit 2mo ago
Labs spend billions hiring experts to generate new data, and better models can better filter existing training data and generate new synthetic data. There’s no reason for that to run out, it’s just expensive.
You could view this as just continually patching a leaky ship. But it seems to work.
- dominotw 2mo ago> Labs spend billions hiring experts to generate new data I thought this is mostly RL data. In my previous comment i was referring to pertaining data.
- sothatsit 2mo agoI remember listening to Andrej Karpathy talk in a podcast about how synthetic data in particular is used to generate more data for pre-training. I see no reasons for that to have changed. I think it is likely a lot of the new data they are paying for contributes to pre-training as well. I would also be very shocked if they weren't filtering or prioritising existing pre-training data as well, for example to do curriculum learning or to avoid data that degrades performance.