3 ms·
No. 1. "Grokking" was shown on 4-digit modular arithmetic with a 1-layer transformer; this article extrapolates it to AGI and a $10B training run with exactly
by Amekedl 3mo ago
No.
1. "Grokking" was shown on 4-digit modular arithmetic with a 1-layer transformer; this article extrapolates it to AGI and a $10B training run with exactly zero intermediate evidence.
2. The "small dataset" is 25 trillion tokens - literally the size of current frontier training sets - but calling it small sounds revolutionary.
3. BabyLM has spent 4 years failing to produce grokking on constrained data; the paper gets a footnote saying "those models were too small," which is unfalsifiable until someone burns $10B.
4. Chain-of-thought is already empirically required for frontier performance - it's expensive, bizarre, and nobody predicted it - yet somehow we're supposed to bet the farm on a phenomenon that has never scaled past arithmetic. We need that data, even if it is just "Actually, ..."
5. If you want to chase "recurrent depth", loop transformers rumored in Mythos/Fable are at least grounded in actual engineering; grokking-at-scale is just vibes ai bro science.
More data is and will always be the answer. Why are all labs distilling from each other?
- gwern 2mo agoFYI, AI-written comments are banned on Hacker News: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html > Don't post generated text or AI-edited text. HN is for conversation between humans.
- Amekedl 2mo ago? How could Grokking be made workable? Mechanistic Interpretability, like using SAEs (EleutherAI did), also methods Heretic employs to decensor, surely you can see some "modding" does indeed achieve things. FYI, I prefer discussion over being mad