5 ms·
Because they exclusively used a model that was about as big as the original GPT-2. Which, I mean, fair enough within these constraints, but it's cited like it'
by FeepingCreature 7mo ago
Because they exclusively used a model that was about as big as the original GPT-2.
Which, I mean, fair enough within these constraints, but it's cited like it's a universal law.
Really all that can be taken away from the study is "we trained a very small model on data generated from it in a particular way, and this was eventually harmful for the model."
Also note that models are nowadays trained on massively self-generated data (task RL post-training) and it seems to significantly improve their performance.