4 ms·
You are not correct about TPUs being drastically better for GPUs at this. If you look at public benchmarks, both have a similar cost per hardware flop ($0.88/hr
by why_only_15 3y ago
You are not correct about TPUs being drastically better for GPUs at this. If you look at public benchmarks, both have a similar cost per hardware flop ($0.88/hr for 312tflops A100 on GCP, $0.97/hr for 275tflops TPUv4) and both achieve similar model flop:hardware flop ratios (40-60%).
An existence proof that GPU mega-clusters are possible is that GPT-4 cost ~$100m over ~3 months, so ~100m a100-hours / (3 months * 30 days/month * 24 hours/day = 2160 hours) = ~45k a100s collaborating, which is the equivalent of ~10 TPUv4 pods on a single training run.
- vl 3y agoI assume/suspect internally Google has v5 already. The thing about TPU clusters that they have hyper-torus optical interconnect between TPUs. This allows for extremely efficient weight updates. To replicate this with A100s you need very custom hardware/software deployment. But to be fair, I don’t know what is latest and greatest available from NVidia or other clouds in this area right now. EDIT: Looks like NVidia has NVSwitch, which provides interconnect for 256 GPUs. Pretty cool!
- AaronFriel 3y agoNVidia acquired Mellanox in 2019 in order to own the interconnect - exactly the custom hardware/software stack you thought needed.
- rerx 3y agoExternal NVSwitch (for more than 8 GPUs) isn't available for purchase yet. But even now you typically buy your A100 or H100 training GPUs in servers where each GPU has an individual 400 Gbps Mellanox networking adapter, then connect them in a sophisticated switched network that provides full throughput between all GPUs of the cluster. The deployment is not that custom, typically you would let Nvidia and your vendors implement their reference design for you: https://www.nvidia.com/en-us/data-center/dgx-superpod/ https://www.nvidia.com/en-us/data-center/dgx-superpod/ Then run an open-source software stack on top of CUDA etc. I don't know how well the TPU hyper-torus interconnect performs, but the networking topology seems to be less general than switched NVLink or InfiniBand.
- vl 3y agoHow is network adapter connected to GPU? They just seat on the same bus? It looks like in TPU v4 cluster each pod with 2 or 4 (?) TPUs has 6 optical interfaces, which directly connect to next pods. I have no idea how they route though this configuration, but my guess most messages are weight updates, which are essentially broadcasts, so it should work out fine with some basic forwarding.
- rerx 3y agoCan't find documentation online yet for this year's H100 systems, but here's a schematic for an A100 server: https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-to-dgxa100.html#system-topology https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-... There 2 GPUs, 2 NICs and 1 NVME are connected via one PCIe switch. The torus topology makes sense for ring algorithms: Allreduce, Allgather, Reducescatter. For purely data parallel training you could put all model replicas into the same ring (although Nvidia also uses hierarchical algorithms that benefit from lower lately). With added model parallelism one will need smaller rings running concurrently. I guess the TPU cluster layout will then put constraints on the most efficient model architectures (as does the network topology of a GPU cluster).