5 ms·
Grokfast: Accelerated Grokking by Amplifying Slow Gradients
- thesz 2y agoGrokFast strongly reminds me of Stochastic Average Gradient Descent: https://www.cs.ubc.ca/~schmidtm/Courses/540-W19/L12.pdf https://www.cs.ubc.ca/~schmidtm/Courses/540-W19/L12.pdf Both use averaging.
- svara 2y agoGrokking is certainly an interesting phenomenon, but have practical applications of it been discovered yet? I remember seeing grokking demonstrated for MNIST (are there any other non synthetic datasets for which it has been shown?), but the authors of that paper had to make the training data smaller and got a test error far below state of the art. I'm very interested in this research, just curious about how practically relevant it is (yet).
- fwlr 2y agoMy gut instinct from reading about the phenomenon says that a “grokked” model of X parameters on Y tokens is not going to outperform an “ungrokked” model with 2X parameters on 2Y tokens - since “grokking” uses the same resources as parameter and token scaling, it’s simply not a competitive scaling mechanism at the moment. It might make sense in some applications where some other hard limit (e.g. memory capacity at inference time) occurs before your resource limit AND you would still see good returns on improvements in quality, but I suspect those are still fairly narrow and/or rare applications.
- joelthelion 2y agoWouldn't it be super useful in cases where data is limited?
- d3m0t3p 2y agoAccording to https://arxiv.org/abs/2405.15071 https://arxiv.org/abs/2405.15071 their grokked model outperformed GPT4 and Gemini1.5 on the reasoning task. We can then argue if the task makes sense and the conclusion stands for other use cases but i think grokking can be useful
- Legend2440 2y agoNobody is really looking for practical applications for it, and you shouldn't necessarily expect them from this kind of academic research.
- svara 2y agoThat doesn't sound right at all. Improving generalization in deep learning is a big deal. The phenomenon is academically interesting either way, but e.g. making sota nets more training data economical seems like a practical result that might be entirely within reach.
- whimsicalism 2y agoi think y'all are both right. grokking is a phenomenon that by definition applies to severely overfit neural networks, which is a very different regime than modern ML - but we might learn something from this that we can use to improve regularization
- barfbagginus 2y agoLooks like grokking could give better reasoning and generalization to LLMs, but I'm not sure how practical it would be to overfit a larger LLM See: https://arxiv.org/abs/2405.15071 https://arxiv.org/abs/2405.15071
- whimsicalism 2y agogrokking doesn't and will not have practical uses, imo - it is just an experiment that revealed cool things that we mostly already suspected about implicit regularization however, techniques we learn from grokking about implicit regularization might be helpful for the training regimes we actually use
- naasking 2y ago> grokking doesn't and will not have practical uses, imo I'm not so sure. Reasoning is the next big hurdle, and grokking and parametric memory seem very effective here. [1] https://arxiv.org/abs/2405.15071 https://arxiv.org/abs/2405.15071
- deleted 2y ago[deleted]
- esafak 2y agoI missed the beginning of the story. Why and when does grokking occur? It seems to be a case of reaching a new basin, casting doubt on the shallow basin hypothesis in over-parameterized neural networks? The last I checked all the extrema in such models were supposed to be good, and easy to reach?
- killerstorm 2y agoIIRC it was observed in a training mode with weight decay. Perhaps a basin with proper generalization is more stable.
- whimsicalism 2y agoi've worked in this field for 6 years and have never heard of the 'shallow basin hypothesis', care to explain more? is it just the idea that there are many good solutions that can be reached in very different parts of parameter space? all that grokking really means is that the 'correct', generalizable solution is often simpler than the overfit 'memorize all the datapoints' solution, so if you apply some sort of regularization to a model that you overfit, the regularization will make the memorized solution unstable and you will eventually tunnel over to the 'correct' solution actual DNNs nowadays are usually not obviously overfit because they are trained on only one epoch
- dontwearitout 2y agoI haven't heard the term "shallow basin hypothesis" but I know what it refers to, these two papers spring to mind for me: 1) Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs https://arxiv.org/abs/1802.10026 https://arxiv.org/abs/1802.10026 2) Visualizing the Loss Landscape of Neural Nets https://arxiv.org/abs/1712.09913 https://arxiv.org/abs/1712.09913 There's also a very interesting body of work on merging trained models, such as by interpolating between points in weight space, which relates to the concept of "basins" of similar solutions. Skim the intro of this if you're interested in learning more: https://arxiv.org/abs/2211.08403 https://arxiv.org/abs/2211.08403
- 2y ago
- buildbot 2y agoWhy only MNIST and a Graph CNN? Those are small and somewhat odd choices. Scale these days should be at least 100 million param models and something like OpenWebText as a dataset in my opinion. Not sure what the SoTA is for visionm but same argument there.
- dzdt 2y agoThis paper is from a small group at an academic institution. They are trying to innovate in the idea space and are probably quite compute constrained. But for proving ideas smaller problems can make easier analysis even leaving aside compute resources. Not all research can jump straight to SOTA applications. It looks quite interesting, and I wouldn't be surprised to see it applied soon to larger problems.
- buildbot 2y agoI’ve been in a small group at an academic institution. With our meager resources we trained larger models than this on many different vision problems. I personally train LLMs on OpenWebText than this using a few 4090s (not work related). Is that too much for a small group? MNIST is solvable using two pixels. It shouldn’t be one of two benchmarks in a paper, again just in my opinion. It’s useful for debugging only.
- all2 2y agoAgain, a small academic institution may not have the experience or know-how to know these things.
- olnluis 2y agoI thought so at first, but the repo's[0] owner and the first name listed in the article has Seoul National University on their Github profile. Far away from a small academic institution. [0]: https://github.com/ironjr/grokfast https://github.com/ironjr/grokfast
- whimsicalism 2y ago
- curious_cat_163 2y agoCute! The signal processing folks have entered the room... :)
- utensil4778 2y agoI'm really annoyed that AI types are just stealing well established vocabulary from everywhere and assigning new arbitrary definitions to them. You have countless LLMs, use one of them to generate new names that don't require ten billion new disambiguation pages in Wikipedia.
- eigenvalue 2y agoI have a suspicion that this technique will prove most valuable for market oriented data sets (like price related time series), where there isn't necessarily that much massive data scale compared to text corpora, and where there are very tight limits on the amount of training data because you only want to include recent data to reduce the chances of market regime changes. This approach seems to shine when you don't quite have enough training data to completely map out the general case, but if you train for long enough naively, you can get lucky and fall into it.
- aoeusnth1 2y agoHow does this differ from momentum in practice? Gradient momentum already applies an exponential-decay average to the gradients. The authors discuss how their approach differs from momentum in formula, but not how it differs in practice. Essentially, momentum and Adam and all other second order optimizers already have explored this intellectual space, so I’m not sure why this paper exists unless it has some practical applications on top of the existing practice.