3 ms·
The grokking papers show that after sufficient training models can transition into a regime where both training and test error gets arbitrarily small. Yes, thi
by smallnamespace 2y ago
The grokking papers show that after sufficient training models can transition into a regime where both training and test error gets arbitrarily small.
Yes, this is out of reach of how we train most models today. But it demonstrates how even current models are capable of building circuits that perfectly predict (meaning understand the actual dynamics) of data given sufficient exposure.
- godelski 2y agoI have some serious reservation about the grokking papers and there's the added complication that test performance is not a great proxy for generalization performance. It is naive to assume the former begets the latter because there are many underlying assumptions there that I think many would not assume are true once you work them out. (Not to mention the common usage of t-SNE style analysis... but that's a whole other discussion) It is important to remember that there are plenty of alternative explanations to why the "sudden increase" in performance happens. I believe if people had a deeper understanding of how metrics work that the phenomena would become less surprising and make one less convinced that scale (of data and/or model) will be insufficient to create general intelligence. But this does take quite a bit of advanced education (that is atypical from a ML PhD) and you're going to struggle to obtain it "in a few weekends".