3 ms·
I just learned about grokking; reminds me of double descent, and I looked up a 2022 paper called "Unifying grokking and double descent". I'm still unclear on wh
by Scene_Cast2 2y ago
I just learned about grokking; reminds me of double descent, and I looked up a 2022 paper called "Unifying grokking and double descent". I'm still unclear on what the difference is. My basic understanding of double descent was that the regularization loss made the model focus on regularization after fitting the train data.
- tehsauce 2y agoGrokking is a sudden huge jump in test accuracy with increasing training steps, well after training accuracy has fully converged. Double descent is test performance increasing, decreasing, and then finally rising again as model parameters are increased.
- scarmig 2y agoWhat they share is a subversion of the naive framework that ML works simply by performing gradient descent over a loss landscape. Double descent subverts it by showing that learning isn't monotonic in parameter count; grokking subverts it by learning after training convergence. I'd put the lottery ticket hypothesis in the same bucket of "things that may happen that don't make sense at all for a simple optimization procedure."
- baq 2y agoMy takeaway from the paper is that you can guide training by adding/switching to a more difficult loss function after you got the basics right. Looks like they never got to overfitting grokking, so maybe there’s more to discover further down the training alley.