3 ms·
Grokking may not even occur for datasets of that scale. Even the MNIST experiments require dropping the training data size from 50k examples to 1k. The reason f
by Imnimo 2y ago
Grokking may not even occur for datasets of that scale. Even the MNIST experiments require dropping the training data size from 50k examples to 1k. The reason for this is that the phenomenon seems to occur at a critical zone of having just barely enough training data to make generalization possible. See https://arxiv.org/abs/2205.10343 https://arxiv.org/abs/2205.10343 for details.
Even figuring out how to induce grokking behavior on a 100M model or OpenWebText would be a big leap in the understanding of grokking. It's perfectly reasonable for a paper like this to show results on the standard tasks for which grokking has already been characterized.