6 ms·
> an IB switch is in the same ballpark as an ethernet switch with the same port speed And both price tags will make Elon's "someone's scamming me with a 'you'r
by creshal 2y ago
> an IB switch is in the same ballpark as an ethernet switch with the same port speed
And both price tags will make Elon's "someone's scamming me with a 'you're an enterprise customer' surcharge" sense tingle. The price tags for anything enterprise networking related are seriously inflated, and I would not be surprised if just making your own NICs and switches is cheaper once you hit a certain deployment size.
- _zoltan_ 2y agoNetworking has never been so cheap at the highest end. Look at the road we've been through in the last 8 in years, rapidly going from 40 to 100 to 400 (200 was somewhat of a dud, 400 came too early) to 800 to 1600Gbps. It's amazing. I'm having trouble feeding things at 400GB/s (not a typo, it's gigabyte/s) per H100 box. For 10 boxes ideally you want 4TB/s...
- starspangled 2y agoNo, but this isn't high end (in 2024), or "enterprise". It's their own designed 100Gb dumb NIC.
- _zoltan_ 2y agoI've replied in a thread, and what I've replied to above was about the enterprise tax and that it's surely cheaper to do your own, not if the original article is about enterprise or not.
- starspangled 2y agoI read the thread. What you replied to was this, > > an IB switch is in the same ballpark as an ethernet switch with the same port speed > And both price tags will make Elon's "someone's scamming me with a 'you're an enterprise customer' surcharge" sense tingle. In the context of Tesla doing their own protocol and not-high-end NICs.
- rv3909i 2y agoIf you were starting a team from scratch this is surely true. But, if you can leverage existing teams and infrastructure, it’s very possible.
- rv3909i 2y agoOn the other hand, silicon has never been so cheap. The hardest things would've been the DDR4 and PCIE interface. But as they're using standard interfaces, and last generation. I'm sure they got a good discount on all that IP and it didn't cost them hardly any man hours to integrate. And Tesla might've even already had the licenses and IP setup as they make other ASICs. I didn't do a budget or anything, but at even 10Ks of units, I could see how this could save money. Or at least not loose money. Assuming a comparable IB network card is ~$1000, which I also didn't price. And there could be other potential cost offsetting features, like power savings.
- KaiserPro 2y ago> The price tags for anything enterprise networking related are seriously inflated I mean yeah, but thats why you have negotiators. List price is what suckers pay. As soon as you start to buy in job lots, or the total price comes to >$500k then stuff becomes a lot cheaper all of a sudden (within reason) Having said that Infiniband is an arse to deploy, but not as much as custom networking protocol on custom silicon.
- zaphirplane 2y agoIdeas you’ll never hear at Google or meta
- vitus 2y ago> > I would not be surprised if just making your own NICs and switches is cheaper once you hit a certain deployment size. > Ideas you’ll never hear at Google or meta You'd be surprised. Google has a very strong tradition of "not-invented here" which extends to some of our production networking gear as well. To be fair, at the time, some of this was justified because the available devices on the market couldn't support our use cases back then. Per section 3.2 of the 2013 B4 paper [0]: Even so, the main reason we chose to build our own hardware was that no existing platform could support an SDN deployment, i.e., one that could export low-level control over switch forwarding behavior. Any extra costs from using custom switch hardware are more than repaid by the efficiency gains available from supporting novel services such as centralized TE. https://cseweb.ucsd.edu/~vahdat/papers/b4-sigcomm13.pdf https://cseweb.ucsd.edu/~vahdat/papers/b4-sigcomm13.pdf
- sangnoir 2y agoYou couldn't have picked a better/worse duo in tech to be wrong on such an assertion. Adding to sibling comment about Google, Meta[1] built 2 large-scale production training clusters for science: one with Infiniband, the other one with a custom RDMA over RoCE fabric. > Custom designing much of our own hardware, software, and network fabrics allows us to optimize the end-to-end experience for our AI researchers while ensuring our data centers operate efficiently. > With this in mind, we built one cluster with a remote direct memory access (RDMA) over converged Ethernet (RoCE) network fabric solution based on the Arista 7800 with Wedge400 and Minipack2 OCP rack switches. Google, Meta and Netflix are among the most obsessive on optimizing their infrastructure - it's bold to assume they haven't looked at their COTS network gear and thought "hmmm..." 1. https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/ https://engineering.fb.com/2024/03/12/data-center-engineerin...
- zaphirplane 2y agoI mean the exact opposite, those were common what ifs