4 ms·
External NVSwitch (for more than 8 GPUs) isn't available for purchase yet. But even now you typically buy your A100 or H100 training GPUs in servers where each
by rerx 3y ago
External NVSwitch (for more than 8 GPUs) isn't available for purchase yet. But even now you typically buy your A100 or H100 training GPUs in servers where each GPU has an individual 400 Gbps Mellanox networking adapter, then connect them in a sophisticated switched network that provides full throughput between all GPUs of the cluster. The deployment is not that custom, typically you would let Nvidia and your vendors implement their reference design for you: https://www.nvidia.com/en-us/data-center/dgx-superpod/ https://www.nvidia.com/en-us/data-center/dgx-superpod/ Then run an open-source software stack on top of CUDA etc.
I don't know how well the TPU hyper-torus interconnect performs, but the networking topology seems to be less general than switched NVLink or InfiniBand.
- vl 3y agoHow is network adapter connected to GPU? They just seat on the same bus? It looks like in TPU v4 cluster each pod with 2 or 4 (?) TPUs has 6 optical interfaces, which directly connect to next pods. I have no idea how they route though this configuration, but my guess most messages are weight updates, which are essentially broadcasts, so it should work out fine with some basic forwarding.
- rerx 3y agoCan't find documentation online yet for this year's H100 systems, but here's a schematic for an A100 server: https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-to-dgxa100.html#system-topology https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-... There 2 GPUs, 2 NICs and 1 NVME are connected via one PCIe switch. The torus topology makes sense for ring algorithms: Allreduce, Allgather, Reducescatter. For purely data parallel training you could put all model replicas into the same ring (although Nvidia also uses hierarchical algorithms that benefit from lower lately). With added model parallelism one will need smaller rings running concurrently. I guess the TPU cluster layout will then put constraints on the most efficient model architectures (as does the network topology of a GPU cluster).