3 ms·
Can't find documentation online yet for this year's H100 systems, but here's a schematic for an A100 server: https://docs.nvidia.com/dgx/dgxa100-user-guide/intr
by rerx 3y ago
Can't find documentation online yet for this year's H100 systems, but here's a schematic for an A100 server: https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-to-dgxa100.html#system-topology https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-...
There 2 GPUs, 2 NICs and 1 NVME are connected via one PCIe switch.
The torus topology makes sense for ring algorithms: Allreduce, Allgather, Reducescatter. For purely data parallel training you could put all model replicas into the same ring (although Nvidia also uses hierarchical algorithms that benefit from lower lately). With added model parallelism one will need smaller rings running concurrently. I guess the TPU cluster layout will then put constraints on the most efficient model architectures (as does the network topology of a GPU cluster).