6 ms·
You would likely be limited by the communication latency between nodes, unless you come up with some unique model architecture or training method. Most of these
by ftufek 4y ago
You would likely be limited by the communication latency between nodes, unless you come up with some unique model architecture or training method. Most of these large scale models are trained on GPUs using very high speed interconnects.