3 ms·
> nanochat is also inspired by modded-nanoGPT Nice synergy here, the lineage is: Karpathy's nano-GPT -> Keller Jordan's modded-nanoGPT (a speedrun of training
by montebicyclelo 1y ago
> nanochat is also inspired by modded-nanoGPT
Nice synergy here, the lineage is: Karpathy's nano-GPT -> Keller Jordan's modded-nanoGPT (a speedrun of training nanoGPT) -> NanoChat
modded-nanoGPT [1] is a great project, well worth checking out, it's all about massively speeding up the training of a small GPT model.
Notably it uses the author's Muon optimizer [2], rather than AdamW, (for the linear layers).
[1] https://github.com/KellerJordan/modded-nanogpt https://github.com/KellerJordan/modded-nanogpt
[2] https://kellerjordan.github.io/posts/muon/ https://kellerjordan.github.io/posts/muon/
- deleted 1y ago[deleted]
- varunneal 1y agoMuon was invented by Keller Jordan (and then optimized by others) for the sake of this speedrunning competition. Even though it was invented less than a year ago, it has already been widely adopted as SOTA for model training
- tbalsam 1y agoThis is the common belief but not quite correct! The Muon update was proposed by Bernstein as the result of a theoretical paper suggesting concrete realizations of the theory, and Keller implemented it and added practical things to get it to work well (input/output AdamW, aggressive coefficients, post-Nesterov, etc). Both share equal credit I feel (also, the paper's co-authors!), both put in a lot of hard work for it, though I tend to bring up Bernstein since he tends to be pretty quiet about it himself. (Source: am experienced speedrunner who's been in these circles for a decent amount of time)
- varunneal 1y agoI think it's good to bring up Bernstein & Newhouse as well as Yuchen Jin, Jiacheng You and the other speedrunners who helped iterate on Muon. But I think it's very fair to call Keller Jordan the main author of Muon of its current form. I'm also in the speedrunning community though maybe not as long as you have
- swyx 1y agosharing some useful resrources for learning Muon (since I'm also just catching up on it) - https://x.com/leloykun/status/1846842883967692926 https://x.com/leloykun/status/1846842883967692926 - https://www.yacinemahdid.com/p/muon-optimizer-explained-to-a-toddler https://www.yacinemahdid.com/p/muon-optimizer-explained-to-a...
- cantor_S_drug 1y agoThis Simple Optimizer Is Revolutionizing How We Train AI [Muon] https://www.youtube.com/watch?v=bO5nvE289ec https://www.youtube.com/watch?v=bO5nvE289ec I found the above video as a good introduction.
- ComplexSystems 1y agoI haven't heard of this before. Has Muon dethroned Adam and AdamW as the standard general purpose optimizer for deep learning?
- spyder 1y agoIt's for hidden layers and not for every parameter: From Keller's Muon github page: "Muon is an optimizer for the hidden weights of a neural network. Other parameters, such as embeddings, classifier heads, and hidden gains/biases should be optimized using standard AdamW." And I just looked into this nanochat repo and it's also how it's used here. https://github.com/karpathy/nanochat/blob/dd6ff9a1cc23b38ce69ddc119fb220f9ee96cedd/scripts/base_train.py#L129 https://github.com/karpathy/nanochat/blob/dd6ff9a1cc23b38ce6...
- kouteiheika 1y agoThe most exciting thing about Muon for me is that it requires half the state of Adam while having either equivalent or better performance. That's amazing if you are VRAM limited! And just like Adam, you can also quantize it. I can get it to work relatively well as low as 4-bit, which essentially cuts down the memory requirements from full 32-bit Adam by a factor of 16x! (And by a factor of 4x vs 8-bit Adam).
- echelon 1y ago8xH100 is pretty wild for a single inference node. Is this what production frontier LLMs are running inference with, or do they consume even more VRAM/compute? At ~$8/hr, assuming a request takes 5 seconds to fulfill, you can service roughly 700ish requests. About $0.01 per request. Is my math wrong?