4 ms·
These seem classic challenges with running distributed systems loads that are not specific to training LLMs. Anyone of the super computers listed here https://
by idkdotcom 2y ago
These seem classic challenges with running distributed systems loads that are not specific to training LLMs.
Anyone of the super computers listed here https://en.wikipedia.org/wiki/TOP500 https://en.wikipedia.org/wiki/TOP500 suffers from the same issues.
Think about it. While the national labs use these systems to model serious stuff -such as climate or nuclear weapons- Meta uses them to train LLMs. What a joke, honestly!
- whiplash451 2y agoA lot of serious things look like a toy or a joke at first.
- idkdotcom 2y ago[dead]
- mhandley 2y agoOn the other hand, Meta just rapidly built two different training networks in existing datacenter buildings, with existing cooling constraints, using mostly commodity components (albeit expensive commodity components) each of which would place at #3 on that top500 list in terms of GPU power. Compare that with how long it took to get any of the other supercomputers from design to being fully commissioned.
- idkdotcom 2y ago[dead]
- _zoltan_ 2y agoFor profit is not less serious than what research labs do. I'd even say it's more important: they drive the economy.
- idkdotcom 2y ago[dead]