4 ms·
Are there any good resources out there describing in practice how existing training workloads are distributed among GPUs? (using tensorflow, pytorch, or whatev
by frogblast 5y ago
Are there any good resources out there describing in practice how existing training workloads are distributed among GPUs? (using tensorflow, pytorch, or whatever else?).
I'm curious how the problem effectively gets sliced.
- singhrac 5y agoSOTA on the biggest language models (which is where effectively the largest models are) is here: https://www.microsoft.com/en-us/research/blog/zero-infinity-and-deepspeed-unlocking-unprecedented-model-scale-for-deep-learning-training/ https://www.microsoft.com/en-us/research/blog/zero-infinity-...