7 ms·
A RoCE network for distributed AI training at scale
- jauntywundrkind 2y agoFrom the paper, seems like they are using RDMA to/from video cards, skipping the nic. > * These transactions require GPU-to-RDMA NIC support for optimal performance* Remarkably consumer computing actually has similarly found reason to bypass sending data through the cpu; texture streaming. DirectStorage and Sony's Kraken purport to let the GPU read direct from the SSD. It's a storage application instead of NIC, but still built around PCIe DMA-P2P (at least the DirectStorage is I think). Table 2, network stats for 128 GPUs is kind of interesting. Most topologies such as AllGather and AllReduce run with only 4 Queue Pairs. Not my area of expertise at all but wow that seems tiny! All this network, and basically everyone's talking to only a few peers? That's what it means right? The discussion at the end of the paper talked about Flowlets. The description makes me think a little bit of hash bucket chaining, where you try the first path, and if latter a conflict arise or the oath degrades, there's a fallback path already planned. Like there's would be a fallback chained bucket in a hash.
- wmf 2y agoThe NIC is still there but they're skipping the data copy from system RAM to GPU RAM. https://developer.nvidia.com/gpudirect https://developer.nvidia.com/gpudirect
- mrlongroots 2y ago> still built around PCIe DMA-P2P Right! GPUs are bandwidth behemoths. If you have a 400 Gbit/s network and a GPU that can saturate it, there's a very good bandwidth reason to cut the CPU out of the path. For collectives, you are probably optimizing for latency vs bandwidth, but in all cases the less you move the data the better it is. > Most topologies such as AllGather and AllReduce run with only 4 Queue Pairs. Not my area of expertise at all but wow that seems tiny! All this network, and basically everyone's talking to only a few peers? That's what it means right? 1. It seems that the cardinality of the collective groups is quite small (~128 GPUs), despite the reported cluster being 16k GPUs. They attribute it to "multidimensional parallelism" and while I don't speak much ML, I'm guessing that it's some sort of horizontal + vertical partitioning, and collectives are only triggered within a partition. 2. As these are collectives, all nodes do participate, but you create an overlay network with a ring or tree topology per collective. In a ring, every node just needs a QP each for left and right neighbors. For trees, it depends on the fanout. Now you could have shallow trees with larger fanouts, as latency is a function of the depth, and I feel like you can have more than 4 QPs (~500 is where NIC SRAM becomes an issue AFAIK). I think there is some buffering overhead per QP (beyond the standard NIC state) that they are trying to minimize, haven't read the paper thoroughly. 3. You do need a fancy network for collectives. They are a latency-sensitive operation on the critical path. The entire multimillion dollar thing is idle while an AllReduce is active, so 500us vs 100us in E2E latency matters a lot.
- latchkey 2y agoBoth Intel (Gaudi 3) and Tenstorrent (Wormhole) have built NIC's into the GPUs. I think that we are going to see more of that in the future, but the only problem with this is that your failure cases become more extreme. NIC dies, I swap a NIC... GPUNIC dies... I swap both. 128 NICs/GPUs is the largest single non-blocking 400G switch size today using a broadcom t5 chip [0]. If you want to go larger than that, you have to go with 6x switches in a spine/leaf configuration, which greatly increases the costs. We're (Hot Aisle) in the process of deploying a cluster of 128 MI300x right now with this configuration using a Dell Z9864F switch [1]. https://hotaisle.xyz/networking/ https://hotaisle.xyz/networking/ and https://hotaisle.xyz/compute/ https://hotaisle.xyz/compute/ [0] https://www.broadcom.com/products/ethernet-connectivity/switching/strataxgs/bcm78900-series https://www.broadcom.com/products/ethernet-connectivity/swit... [1] https://www.delltechnologies.com/asset/en-us/products/networking/technical-support/dell-powerswitch-z9864f-on-spec-sheet.pdf https://www.delltechnologies.com/asset/en-us/products/networ...
- eslaught 2y agoSo they're re-inventing HPC networks in the data center. https://en.wikipedia.org/wiki/Fat_tree https://en.wikipedia.org/wiki/Fat_tree https://www.cs.umd.edu/class/spring2021/cmsc714/readings/Kim-Dragonfly.pdf https://www.cs.umd.edu/class/spring2021/cmsc714/readings/Kim... I'm sure there are innovations here, but most of this has been standard in HPC for decades. (Fat trees since 1985, Dragonfly since 2008.) This is not new science, folks.
- wmf 2y agoIt's not new science, but tuning RoCE performance is new engineering.
- mrlongroots 2y agoAnd honestly the difference between science and engineering in systems is overrated. The real difference is "systems that work, or everyone knows how to run" vs "systems that are a bit of black magic". While on paper it is true that HPC has had collectives and dragonflies and RDMA for ages, a whole host of issues make tuning these things hard and tedious. Such work also informs the design of future "smart networks" by keeping everyone updated with common problems and workarounds that folks employ.
- teleforce 2y agoInteresting approach on distributed AI training albeit a very expensive one. Personally I'm baffled why no one has come up with a similar project to SETI@home or Great Internet Mersenne Prime Search in harnessing truly distributed and low cost solutions to open model of AI training at scale [1],[2]. [1] SETI@home: https://setiathome.berkeley.edu/ https://setiathome.berkeley.edu/ [2] Great Internet Mersenne Prime Search: https://en.wikipedia.org/wiki/Great_Internet_Mersenne_Prime_Search https://en.wikipedia.org/wiki/Great_Internet_Mersenne_Prime_...
- immibis 2y agoUnlike those projects, AI training requires very high bandwidth between nodes.
- sooperserieous 2y agoThey have: http://salad.com http://salad.com and others
- dan-robertson 2y agoI think there’s two reasons: 1. AI training for large models just requires a lot of bytes – there are lots of parameters and lots of training data, and I think you don’t provide such useful training by just reusing the same subset of the training data. Similarly you can’t really train just a fraction of the parameters. The more diverged participants are, the less useful their contributions may be so I think you’d need to move the parameters back and forth quite frequently too 2. People just don’t have that much compute at home. And you probably need more compute because of the bandwidth/latency issues.
- teleforce 2y agoBoth your points are relevant but neither are critical nor crucial for successful training of GPT foundation models. For the first point, unlike real-time queries of ChatGPT prompts, training does not relied on minimum latency for feedback. For the second point, our average laptops and PCs nowadays are much better than the highest end graphics workstation in gaming industry AAA studios about 30 years ago circa 1990s. Last year Google made an interesting memo mentioning that there's no company has the moat on GPT and eventually nobody can compete with open source AI model. Perhaps Google is referring and alluding to the potential of distributed global collaborations for training of foundation models not unlike the SETI@home and similar approaches.
- zuckerma 2y agoThis is slick
- marob 2y ago[flagged]