44 ms·
Yes, ultimately everyone is currently doing something which looks like synchronous data parallel training on the outside. The linked PDF is very light on detai
by adw 2y ago
Yes, ultimately everyone is currently doing something which looks like synchronous data parallel training on the outside.
The linked PDF is very light on detail, but what results they do claim are about a 1.2bn parameter model. This is tiny; you don't need network-bound distributed training (ie, anything beyond a single datacenter class machine, or less if you're patient) to train a model that size. The comms requirements also scale with the model size, so I strongly suspect people hoping for embarrassingly-parallel-style scaling properties are going to be disappointed.
(They also appear to have, in part, reinvented parameter servers.)
- huac 2y agoin particular it appears that they only implement data parallel DP - at 1.2B you can fit full copy of model into memory, but larger models require splitting the weights across multiple machines (different techniques eg distributed data parallel DDP, tensor parallel TP, pipeline parallel TP, ...) without more details it's unclear if the proposed technique keeps its speedups in that case
- itkovian_ 2y agoThis is not true