12 ms·
So you want to rent an NVIDIA H100 cluster? 2024 Consumer Guide
- latchkey 2y agoGreat post. The ethernet section is especially interesting to me. I'm building a cluster of 16x Dell XE9680's (128 AMD MI300x GPUs) [0], with 8x 2p200G broadcom cards (running at 400G), all connected to a single Dell PowerSwitch Z9864F-ON, which should prevent any slowness. It will be connected over rocev2 [1]. We're going with ethernet because we believe in open standards, and few talk about the fact that the lead time on IB was last quoted to me at 50+ weeks. As kind of mentioned in the article, if you can't even deploy a cluster the speed of the network means less and less. I can't wait to do some benchmarking on the system to see if we run into similar issues or not. Thankfully, we have a great Dell partnership, with full support, so I believe that we are well covered in terms of any potential issues. Our datacenter is 100% green and low PUE and we are very proud of that as well. Hope to announce which one soon. [0] https://hotaisle.xyz/compute/ [1] https://hotaisle.xyz/networking/
- omneity 2y agoDid you procure the servers directly from Dell or through a distribution partner?
- latchkey 2y agoKind of both. We started talking to Dell first and then they introduced us to Advizex. We are now effectively partnered with both companies, which is fantastic as the Advizex team have ex-Dell people working directly with us. We are lucky to have two whole teams of super talented people helping us out on this journey.
- logicchains 2y agoMeta had success building an ethernet cluster on Arista 7800 with Wedge400 and Minipack2 OCP rack switches.: https://www.datacenterdynamics.com/en/news/meta-reveals-details-of-two-new-24k-gpu-ai-clusters/ https://www.datacenterdynamics.com/en/news/meta-reveals-deta...
- latchkey 2y agoI actually have a call with Arista next week to learn more about their solutions. Especially those 7800's. They look awesome. One "problem" we have right now is that our cluster cannot support more than 128 GPUs. If we wanted to scale with Dell, we'd have to buy 6x more Z9864F to add one more cluster, which is crazy expensive and complicated. I want to see if Arista has something that can help us. That said, I also have to find a customer that wants more than 128 MI300x and that hasn't happened... yet.
- csmpltn 2y ago> "Our datacenter is 100% green" Cool, where can I read more about this? How do you power your DC?
- walrus01 2y agoPlenty of datacenters that are somewhere generally in the pacific NW can claim to be "Green" because their power supply is entirely hydroelectric. https://www.nwd.usace.army.mil/CRWM/CR-Dams/ https://www.nwd.usace.army.mil/CRWM/CR-Dams/ Many of those areas also happen to have the lowest $ per kWh electricity in North American, the only lower rate is available near a few hydroelectric dams in Quebec.
- latchkey 2y agoPreviously, I had multiple data centers in Quincy, WA. Those were hydro green. It is an area that hosts a whole multitude of big hyperscaler companies.
- mulmen 2y agoAs anyone who has driven through the Columbia River Basin can tell you wind power is also abundant in Washington. The grid is very clean here but it’s certainly not purely hydro.
- dlkf 2y agoWhy is green in scare-quotes?
- uncertainrhymes 2y agoPeople make the argument that is a giant datacenter is consuming 50% of some local hydro installation, everyone else is town is buying something else that is less green. It opens up questions about grids and market efficiency, so your mileage may vary.
- sangnoir 2y ago> People make the argument that is a giant datacenter is consuming 50% of some local hydro installation, everyone else is town is buying something else that is less green. I don't think that's a cogent argument. It's akin to saying a vegan commune in a small is buying is buying up 50% of the vegan food, "forcing" others to buy meat-products, and framing this to cast doubts on whether they are truly vegan. Consumers aren't in a position to solve supply problems.
- dpflan 2y agoWhat are you using these for? Providing compute for customers that want to train/infer? What is the level of interest and level of success customers are seeing using these services Hot Aisle offers?
- latchkey 2y agoWe are a bare metal compute offering. They can be used for whatever people want to use them for (within legal limits, of course). Interest is much higher now that we've started to work with teams publishing benchmarks which show that H100's have a nice competitor [0]. I'll admit, it is still early days. We just finished up another free compute [1] two week stint with a benchmarking team. One thing we discovered is that saving checkpoints is slow AF. I'm guessing an issue with ROCm. Hopefully get that resolved soon. Now we are in the process of onboarding the next team. [0] https://hotaisle.xyz/benchmarks-and-analysis/ [1] https://hotaisle.xyz/free-compute-offer/
- rvnx 2y ago> We would love to offer hourly on-demand rates for individual GPUs, but we can't do so at this time due to a limitation in the ROCm/AMD drivers. This limitation prevents PCIe pass-through to a virtual machine, making multi-tenancy impossible. AMD is aware of this issue and has committed to resolving it. One idea to help you: Are you sure you need a virtual machine ? Couldn't you boot the machines under PXE to solve the imaging problem ? Essentially you have TFTP server that gives a Linux image and boot on it directly
- latchkey 2y ago1 chassis, 8 gpus. We want to be able to break that chassis up into individual GPUs and allocate 1 GPU to 1 "machine". I previously PXE booted 20,000 individual playstation 5 diskless blades and I'm not sure how PXE would solve this. The only alternative right now is to do what runpod (and AMD's aac) are doing and do docker containers. But that has the limitation of docker in docker, so people end up having to repackage everything. You also can't easily run different ROCm versions since that comes from the host, and if you have 8 people on a single chassis... it becomes a nightmare to manage it. We're just patiently waiting for AMD to fix the problem.
- deleted 2y ago[deleted]
- rlupi 2y agohttps://docs.nvidia.com/dgx-superpod/reference-architecture-scalable-infrastructure-h100/latest/network-fabrics.html https://docs.nvidia.com/dgx-superpod/reference-architecture-... NVIDIA large GPU supercomputers have separate compute-networking (between GPUs) and storage-networking (storage to GPUs, or storage to SSD, SSD to GPUs with CPU assistance). This helps avoid networking issues, even more if not using Infiniband. From what I read here and on your website, you don't go that route. I haven't found the equivalent system level reference architecture for MI300x from AMD. I wonder if you have a link to a public document where AMD provides guidance about this choice?
- derefr 2y ago> even more if not using Infiniband It's interesting that the above "HPC reference architecture" shows a GPU-to-GPU Infiniband fabric, despite Nvidia also nominally pushing NVLink Switch (https://www.nvidia.com/en-us/data-center/nvlink/ https://www.nvidia.com/en-us/data-center/nvlink/) for the HPC use-case.
- bee_rider 2y agoHow does NVLink work? Because I already know MPI and I’m not going to learn anything else, lol. Edit: after googling it looks like OpenMPI has some NVLink support, so maybe it is OK.
- zxexz 2y agoI use OpenMPI with no issues over multiple H100 nodes and A100 nodes, with multiple infiniband 200G and ethernet 100G/200G networks, and RDMA (though using mellanox instead of broadcom cards, but afaik broadcom supports this just the same). Side note, make sure you compile nvidia_peermem correctly if you want GDRMA to work :)
- latchkey 2y agoNo issues, except this minor bit of arcane knowledge that is missing from SO. :)
- RobRivera 2y agoNice! Whats benchmark standard these days? Still Superbench or yall have something inhouse? Re: lead time quote :O I guess I got spoiled working for one of the major cloud vendors. The thought of poor b2b vendor support never entered my risk matrix. If you own your own cluster, the network bottleneck becomes less a dollar cost I suppose, since you arent being charged a premium to rent someone elses compute
- latchkey 2y agoWe don't run benchmarks ourselves. We donate the expensive Ferrari worth of compute for others to do it. This is the most unbiased way I could think of getting useful real world data to share. The 3rd team just finished up and the 4th is getting started now. I've got 23 others in the wings. https://hotaisle.xyz/free-compute-offer/ https://hotaisle.xyz/free-compute-offer/
- pheatherlite 2y agoBeing out of the loop for awhile. Has amd made anything similar to cuda? Are cots frameworks such as pytorch and tensorflow on par when running on amd hardware? What makes investing in amd cluster/chips worthwhile?
- jononor 2y agoROCm is the framework. Both PyTorch and Tensorflow have versions that support it. Have no experience with it, so cannot say how it works in practice.
- JackYoustra 2y agoWhy xeon instead of epyc?
- latchkey 2y agoSadly, the only option Dell supports today. To get to market the fastest, they took their existing H100 chassis solution, swapped out the GPUs/baseboard and called it a day.
- startupsfail 2y agoIt seems the reliability, speed and scalability drops with the Ethernet are somewhat manageable. To quote the article - From our tests, we found that Infiniband was systematically outperforming Ethernet interconnects in terms of speed. When using 16 nodes / 128 GPUs, the difference varied from 3% to 10% in terms of distributed training throughput[1]. The gap was widening as we were adding more nodes: Infiniband was scaling almost linearly, while other interconnects scaled less efficiently. And then they do mention that the research team needs to debug unexplained failures on Ethernet that they’ve not seen on Infiniband. This actually can be the expensive part. Particularly if the failures are silent and cause numerical errors only.
- latchkey 2y agoA single switch should mitigate some of the throughput issues. As for issues, this is why I have a full professional support contract with Dell and Advizex. If there are issues in the gear, they will step in to help out. Especially on the switch, since it is a spof, we went with a 4 hour window.
- teaearlgraycold 2y agoWe just set up a small cluster of our own. We’re not using infiniband but it didn’t seem like it would be a 50 week lead time to setting it up. Where did you get that number?
- latchkey 2y agoI hear a couple issues with your comment... "small cluster", "we're not using infinband" I'm only saying what was quoted to me. The cards are easy to get, it is the switches that are more difficult.
- teaearlgraycold 2y agoI saw plenty of used switches available. Are current gen switches necessary?
- latchkey 2y agohttps://news.ycombinator.com/item?id=40951133 https://news.ycombinator.com/item?id=40951133
- teaearlgraycold 2y agoSounds reasonable. We’re in a very different situation. We bought a bunch of used hardware but our servers are for training models. All of our production stuff runs in AWS. We definitely don’t want downtime but a day where researchers have issues is very different than a day where your app is down and your reputation is tarnished.
- latchkey 2y agoYour customer is yourself. My customer is hopefully you. I do not want to piss you off, especially by telling you that the failure is because I bought used unsupported hardware.
- zxexz 2y agoIf you’re OK with used equipment, I find Infiniband basically on par pricewise with used ethernet gear. Mellanox stuff is super easy to test and you can run everything with a mainline kernel. MQM8700/MQM8790 switches can be had really cheap now used. Cables are dumb expensive...except a dime a dozen used. Or buy new, 100% compatible from FS or the likes. ConnectX-6 cards are quite cheap used thanks to HFT firms constantly upgrading to the latest, and many can be set to Infiniband OR Ethernet (I _think_ some of the dual port cards suppirt simultaneous). I set up a cluster that right now has like 70 of these, at least half of them used. Have not had an issue with any of them yet (I botched a couple firmware updates but every time I saw able to just reset the card and push firmware again). Every machine is connected to the IB fabric AND to at least 100G ethernet. FWIW, used 100G Ethernet equipment is now cheap enough I’ve been upgrading my home network to be 100G. Cheaper than new consumer 10G equipment.
- latchkey 2y ago> If you’re OK with used equipment I'm building a business, not a home lab. (before you continue to downvote me, read what I wrote below)
- adastra22 2y agoSo? The difference in cost can be as much as 10x. If you're building a startup, that matters.
- latchkey 2y agoWhat matters is support contracts and uptime. If you have a dozen customers on a server that cannot access things because of an issue, then as a startup, without a whole customer support department, you're literally screwed. I've been on HN long enough to have seen plenty of companies get complaints after growing too quickly and not being able to handle the issues they run into. I'm building this business in a way to de-risk things as much as possible. From getting the best equipment I can buy today, to support contracts, to the best data center to just scaling with revenue growth. This isn't a cost issue, it is a long term viability issue. Home lab... certainly cut as many corners as you want. Cloud service provider building top super computers for rent... not so much. There is a reason why not a lot of people start to do this... it is extremely capital intensive. That is a huge moat and getting the relationships and funding to do what I'm doing isn't easy and took me over 5 years to get to this point of just getting started. I'm not going to blow it all on cutting corners on some used equipment.
- huqedato 2y agoThe only essential aspect this article doesn't answer: How much does it cost? All the rest is metadata. I would have preferred a clear table with vendors, prices and features. And less bla-bla.
- ea016 2y agoI couldn't share any pricing data since the discussions with providers are private. Instead, I added a graph of prices from gpulist.ai. For an Infiniband cluster, median is $2.3 per H100 hour, average is $2.47.
- ilaksh 2y ago$2.47 * 256 * 24 * 30 = $455k ?
- dijit 2y agoBased solely on my own calculations that I made for the board of my company; this is within the parameters I would expect, yeah.
- doesnotexist 2y agoI'm surprised that ~$10 million dollars of GPUs, @ $40k per H100 and excluding operational costs like the energy bill, only rents for $455k per month. Sounds like a really tough business since the amount of time required to recoup the costs of ownership (~21 months) seems like a really long time. A new generation or two of chips will have hit the market in that time, depreciating the recurring rental income. Leads me to wonder how much if any profit can be made renting GPUs.
- silverlake 2y agoGood info! I use an HPC with SLURM. 40k GPUs shared by hundreds of users. It works well enough. I don’t know how the market for cloud-based clusters works. Why didn’t OP use AWS or Google for on-demand training? Is it just down to cost?
- macksd 2y agoIf you do, in fact, need H100s, they can be very hard to get. Even the smaller flavors of A100 you sometimes request, wait days for, and then 1 node might show up during a weekend. And for the reasons described in the article and the fact that large training jobs can be network-limited, nicer networks can be a big deal.
- eigenvalue 2y agoLots of good and detailed information here, thanks. I'm curious why Ethernet interconnect is so unreliable in practice compared to the Infiniband. I would think that at this point, after a decade or more of current Ethernet standards, all the kinks would be worked out and the worst that would happen would be occasional latency spikes and a few lost packets that could be retransmitted quickly. Shouldn't the training frameworks be more robust to that sort of thing?
- rlupi 2y agoInfiniband and ethernet are very different at the lowest levels. Ethernet interconnects use RoCE (RDMA over converged ethernet), which actually encapsulated infiniband transport in ethernet, but you still pay for higher routing latency, and you need separate compute-network and storage-network to avoid queueing (lossless ethernet). https://community.fs.com/article/infiniband-vs-ethernet-which-is-right-for-your-data-center-network.html https://community.fs.com/article/infiniband-vs-ethernet-whic... https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet Also... don't underestimate the PCI bus bottlenecks when you put 8x 400GB networking + 8x GPUs. There are ways now to have a tree of PCI switches and avoid overloading the main one, each GPU gets its own networking card and PCI switch.
- latchkey 2y agoThis is a great comment. Our cluster is 128 GPUs into a single Dell switch... should help with the queuing. We also have a separate e-w 100G network. This is why we went with Dell XE9680 chassis... people forget that PCI switches are quite important with this level of compute. Dell has done a good job here.
- eigenvalue 2y agoInteresting, thanks. From the wikipedia link, this seems like the probable culprit for why things break: "Although in general the delivery order of UDP packets is not guaranteed, the RoCEv2 specification requires that packets with the same UDP source port and the same destination address must not be reordered."
- ec109685 2y agoHow do the large clouds compare from an availability and cost perspective compared to finding a smaller provider and renting a dedicated cluster?
- choppaface 2y agoThe large clouds will often be able to give the biggest players a big discount. Perhaps not on raw GPU prices (maybe extended trial), but if you already have e.g. 1PB in object storage then they might give you 20-30% discount. Moreover if your contract is jumbo, they’ll give you not only a Slack channel but send Forward Deployed Engineers to your office and/or help you build part of the model training software. But for a deployment the size of the OP (Photoroom) I doubt any of the big clouds would offer a discount. Especially if they were not already negotiating with multiple clouds. Probably the best argument for going with a large cloud provider on a smaller budget is that you already use some of their other services significantly and your MLE-to-devops headcount makes something like Photoroom’s test infeasible.
- cavisne 2y agoLarge clouds tend to be not very good for GPU clusters. The security & management of multi-tenant GPU's is very complex. They have to buy from Nvidia so there is no negotiating leverage. And they know you will be a high maintenance customer (wanting all nodes to work at all times). So there is no big price or availability advantage for a large cloud (unless you are large enough to rent a dedicated cluster from a large cloud)
- barbazoo 2y ago> Electricity sources and CO2 emissions I love that they included this in their consideration and pointed out the impact running these GPUs has on the environment.
- storyinmemo 2y agoShameless employer promotion here while I work on H100 clusters today: https://www.datacenterdynamics.com/en/news/crusoe-puts-cpus-into-at-north-data-center-in-iceland/ https://www.datacenterdynamics.com/en/news/crusoe-puts-cpus-..., https://crusoe.ai/cloud/ https://crusoe.ai/cloud/ Iceland seems to have an excess of energy to population and it's very green.
- barbazoo 2y agoI applaud the vision, I wish you folks hired remote software devs
- irq 2y agoCrusoe is unabashedly anti-remote work, which is curious considering their company’s environmental and energy locality focus. I interviewed with them, they are very much an old school “you must all work physically in San Francisco” company. I work for one of their competitors now, one that embraces remote work.
- tucnak 2y agoI'm sure they don't even realise What and Whom they have lost. You've taught them a lesson!!
- ai4ever 2y agoso, ex bitcoin miners are pivoting into gpu clouds whats the point ?
- 123yawaworht456 2y agoit has fuck all any impact if the electricity is sourced from nuclear/hydro/solar/wind/geothermal.
- Jun8 2y agoSay you want to burn about $500 as a curiosity project for 8 nodes for a day. Any suggestions for what job to run?
- teaearlgraycold 2y agoIf you’re just burning money you might as well mine crypto.
- turtles3 2y agoAs a random thought, this seems to be about the same order of magnitude compute as Karpathy's recent GPT-2 work: https://github.com/karpathy/llm.c/discussions/677 https://github.com/karpathy/llm.c/discussions/677 You could take the final checkpoint from that page and run it for some additional steps and see if it improves? You could always publish the final checkpoint and training curves - someone might find it useful.
- fragmede 2y agoYou could benchmark how fast you can count the number of words (and characters and lines) in all of project gutenberg with wc-gpu. https://github.com/fragmede/wc-gpu https://github.com/fragmede/wc-gpu
- 8organicbits 2y agoPretty sparse on pricing data, I guess everyone asked them to keep it private.
- latchkey 2y agoCompute pricing isn't really private. When you're talking about high end compute, pricing is very much case by case. What is the point of posting it if changes due to everyone having different needs? We have base pricing on our website, but I guarantee that if someone comes to me asking for a year reservation, I'm not going to give the quoted price. What I have there is just a good starting point to get the discussion going. I also had a great dialog with GC on LI over their version of this posting, it seems they really value this customer and it is long term relationship. My assumption is the actual pricing reflects that. One other thing on the special needs, GC mentioned they had 3 extra spare chassis in play as uptime was critical. That is not an insignificant amount of investment to have just laying around.
- spott 2y agohttps://gpulist.ai https://gpulist.ai No idea about how accurate that is, but if you want cluster pricing...
- latchkey 2y agoI got them to add the Verified badge, but the rest is pretty much about as accurate as craigslist.
- Der_Einzige 2y agoI hate how all high end markets wage wars on price discovery. Most sell their products through middle-men instead of directly for a reason. High end furniture? Suddenly prices go away and you have to "get quotes". High end GPUs? Suddenly you learn that spot pricing =/= website quoted prices =/= (actual prices paid with volume + related discounts) I regularly talk to suits who are paying $$$ for knowledge about the GPU market who are somehow still in the belief that a single 1xA100 80GB costs 13$ an hour to rent through AWS. When I tried to correct them, they almost seemed not to believe me. Things that take us tech bros minutes (checking the up-to-date price data by going to the screen in your cloud console to spin one up) or hours (emailing your cloud rep for pricing data with discounts) take suits years to poorly approximate knowledge of. If I go into a market, and price discovery isn't easy on a product, I know I'm dancing with a good chance of being scammed, and by the most greedy, comic-book evil kind of rich people. Suits aren't immune to this, and I'm certainly not either.