4 ms·
It was actually news to me recently that you could exploit massive parallelism for deep learning. Doesn't that require fancy networks? Or are you actually able
by TallGuyShort 6y ago
It was actually news to me recently that you could exploit massive parallelism for deep learning. Doesn't that require fancy networks? Or are you actually able to do a meaningful amount of work independently and synchronize less often?
- atalwalkar 6y agoThere are different ways to exploit massive parallelism. One is to perform distributed training of an individual model, and the degree to which you can see good scaling here indeed depends on the network itself (as well as the quality of the distributed training algorithm itself, with things like Horovod being best-in-breed). A second way to exploit parallelism is via hyperparameter search, and in particular when leveraging early-stopping based approaches like Hyperband/ASHA. Speedups in this setting tend be fairly robust to the the choice of network. As you mentioned, in such settings you can do meaningful amounts of work independently, and moreover can leverage asynchrony to further avoid bottlenecks. Disclaimer: I am one of the co-founders of Determined, and also one of the developers of Hyperband/Asha.
- vaidhy 6y agoThere are a few things that people getting started on ML do not know - deep networks are heavily dependent on random initialization - for production scale, it is generally a good idea to have multiple instances of the network initialized and use early-stopping as a way of weeding out poorly initialized network (which can be done in a nice parallel way). You can do either data or model parallelism - TensorFlow 2.x has nice utilities for data parallel. DeepSpeed from Microsoft helps with model parallel training. Slow parameter sync between nodes is a bottleneck.