3 ms·
I think the broader context here is that it's "fine for training" in the sense that you can successfully train a model, but it's not "fine for training" in the
by smallnamespace 3y ago
I think the broader context here is that it's "fine for training" in the sense that you can successfully train a model, but it's not "fine for training" in the sense that it can only train small models due to the lack of scaling across cards, which directly cuts against where ML has been trending over the last several years.
In LLM-land we've rapidly gone from training bespoke models to doing fine tuning to RLHF to zero-shot prompting. The better the underlying model the more you can do without additional training, so hardware that fails to scale up to the largest training runs will have limited practical utility despite technically supporting training.
- buildbot 3y agoWormhole does support scale out though via it's built in networking? And there seem to be links for something on top of the board, NVLINK style. And yes, I know, I've been working on LLM pretraining for about 4 years now, since 2020. The number formats themselves so far are mostly scale invariant or improve with larger scale - you can quantize a larger model and see less performance drop than a smaller model.