6 ms·
This kind of computing must be a different kind of world than the one I work in. 80 microseconds of latency seems high to me when infiniband can do single digi
by cout 2y ago
This kind of computing must be a different kind of world than the one I work in. 80 microseconds of latency seems high to me when infiniband can do single digit latency with unreliable datagrams, which turn out to be mostly reliable due to the credit system.
- exabrial 2y agowhats the credit system?
- xtacy 2y agoOP is referring to "Credit based flow control", which is a way to ensure a sender does not overwhelm a receiver with more data than it can handle. Usually, this is line-rate, but if the other side is slow for whatever reason (say the consumer is not draining data), you wouldn't want the sender to continue sending data. If you also have N hosts sending data to 1 host, you would need some way of distributing the bandwidth among the N hosts. That's another scenario where the credit system comes. Think of it as an admission control for packets so as to guarantee that no packets are lost. Congestion control is a looser form of admission control that tolerates lossy networks, by retransmitting packets should they be lost.
- yumraj 2y ago> infiniband Or I guess even RoCE
- kamikaz1k 2y agoNot really familiar with this space but I think the entire Dojo/DIY strategy was kicked off because Elon wanted to not get cornered on supply or cost by nvidia. And infiniband is an nvidia technology, so they wouldn’t use that simply from strategic POV. Are there other technologies they could have used? Also, the 80us is supposed to be the worst case, where typical is supposed to be <10us. Again not knowing anything about infiniband, what’s the typical perf? I tried to google but the people who are talking about it are in the know in ways I’m not. Thanks!
- justahuman74 2y agoIs RoCE no good?
- publicmail 2y agoThe problem is that it kind of relies on a lossless layer 2 (flow control) which has its own set of problems in large scale networks. This is what things like this try to solve: https://cloud.google.com/blog/topics/systems/introducing-falcon-a-reliable-low-latency-hardware-transport https://cloud.google.com/blog/topics/systems/introducing-fal...
- charleshn 2y agoIndeed, it seems that 80usec is just given as an upper bound based on the 1MB buffer at 100Gbps. It is definitely possible to go much lower than 80usec on Ethernet. But obviously it depends on the scale, utilisation etc. At the sizes of GPU clusters we're talking about these days - 32K and up - things get tricky. The main alternative to Infiniband used in the industry is RoCE - Meta has written a lot about it [0]. There's several reasons to avoid Infiniband, such as cost, availability, vendor lock in, lack of experience etc. Those are some of the reasons why many players are trying hard to make Ethernet work, see Ultra Ethernet [1]. [0] https://engineering.fb.com/2024/08/05/data-center-engineering/roce-network-distributed-ai-training-at-scale/ https://engineering.fb.com/2024/08/05/data-center-engineerin... [1] https://ultraethernet.org/ https://ultraethernet.org/
- 2y ago
- deleted 2y ago[deleted]
- namibj 2y agoAlso PCIe is worth mentioning for it's credit-based full reliability (in the absence of hardware failures, which are still signaled).
- iforgotpassword 2y agoI assume infiniband is much more expensive, but then again you have to offset all the development cost first.
- jabl 2y agoIt's not, really. It's been a while since I've checked pricing so my data might be old, but an IB switch is in the same ballpark as an ethernet switch with the same port speed. Same for HCA's. There's no analogue in the Infiniband world to dirt cheap 1GbE RJ-45 switches though.
- creshal 2y ago> an IB switch is in the same ballpark as an ethernet switch with the same port speed And both price tags will make Elon's "someone's scamming me with a 'you're an enterprise customer' surcharge" sense tingle. The price tags for anything enterprise networking related are seriously inflated, and I would not be surprised if just making your own NICs and switches is cheaper once you hit a certain deployment size.
- _zoltan_ 2y agoNetworking has never been so cheap at the highest end. Look at the road we've been through in the last 8 in years, rapidly going from 40 to 100 to 400 (200 was somewhat of a dud, 400 came too early) to 800 to 1600Gbps. It's amazing. I'm having trouble feeding things at 400GB/s (not a typo, it's gigabyte/s) per H100 box. For 10 boxes ideally you want 4TB/s...
- starspangled 2y agoNo, but this isn't high end (in 2024), or "enterprise". It's their own designed 100Gb dumb NIC.
- _zoltan_ 2y agoI've replied in a thread, and what I've replied to above was about the enterprise tax and that it's surely cheaper to do your own, not if the original article is about enterprise or not.
- starspangled 2y agoAFAIKS the protocol can tolerate up to about 80 microseconds of latency. The graph at the end shows they measured (one way) latency at 1.3 microseconds (compared with 2.0 for IB).
- eecc 2y agoWell, whether it matters depends on the workload: IB is basically remote DMA so if you need to pick and poke remote data I guess it'll work as another NUMA tier. But for AI training, where you're simply shuffling around large stacks of matrices, my guess is latency constraints weaken.