5 ms·
Nice compiled list of stats. I'm not sure what they mean by H100s requiring pre-approval on LambdaLabs? Maybe I was grandfathered in since I had an account pr
by fpgaminer 3y ago
Nice compiled list of stats. I'm not sure what they mean by H100s requiring pre-approval on LambdaLabs? Maybe I was grandfathered in since I had an account prior to their rollout, but I never had to do anything special to rent H100s there.
Somewhat related and hopefully helpful: My experience so far using the H100 PCIes to train ViTs:
Pros: A little over 2x performance compared to A100 40GB. 80GB by default. fp8 support, which I haven't played with but supposedly is another 2x performance win for LLMs. For datacenters and local workstations it's twice as power efficient as the A100s. Supposedly they have better multi-GPU and multi-node bandwidth, but my workload doesn't stress that and I can only rent 1xH100s at the moment.
Cons: I had trouble using them with anything but the most recent nVidia docker containers. The somewhat official PyTorch containers didn't work, nor did brewing my own. Luckily the nVidia containers have worked fine so far. I just don't like that they use nightly PyTorch. In addition to that, because of their increased performance, they really start to push the limits on feeding data fast enough to them with existing system configurations. I was CPU-limited on the LambdaLabs 1xH100 machines because of this.
Overall the pricing has worked out equal for my use-case, but I'm sure fp8 would make them more affordable. If fp8 were available out-of-the-box on PyTorch I'd play with it, but it's only available right now from some nVidia codebase specific to LLMs.
Even at equal pricing, having twice the power per GPU and per node is a big win. That increases experiment iteration across the board.
Side note: If I recall correctly, nVidia is heavily differentiating the H100 products by their interface this go around, which I find quite odd. The SXM version of the cards are supposed to be something like twice as beefy as the PCIE version? Not sure why they're doing that; gonna make comparing rentable instances all the more difficult if you overlook that little detail. A100 had a little bit of this, but the difference was never much in practice except between specifically the A100 40 GB PCIe and the 80 GB SXM, which was something like 10% faster.
Anyone else having fun with the new toy our overlords have allowed us to play with?
- akiselev 3y ago> The SXM version of the cards are supposed to be something like twice as beefy as the PCIE version? Not sure why they're doing that I can't help but feel it's another cloud/enterprise cash grab. The SXM baseboards are way more expensive than server motherboards and NVIDIA makes them.
- buildbot 3y agoIt’s really not (well, no more than PCIE) - SXM has nvlink integrated, and more power delivery built in. Thus, they can crank the max wattage way higher, though it definitely gets into the diminishing return zone quickly past 300W. SXM baseboards are made by more than Nvidia, Dell, HP, and Supermicro all have their own designs. (Disclaimer - I work for MS Azure but have no internal knowledge about the costs/designs/capex whatever of these systems)
- bombcar 3y agoMany cloud providers limit what a brand new account can do. Some automatically drop those limits at a certain age, others require you talk to them.
- sashank_1509 3y agoIf you don't mind answering, what's your use case. Often I feel you either need to do something large scale like a 100 A100s, or you're better off with 8 3090s to run many experiments in parallel
- bluedino 3y agoWe have a couple GPU nodes in our cluster, each have 4 80GB cards (or V100 nodess are only 32GB). We have users that have a workstation with a 3090, but the larger amount of memory is what they are after.
- fpgaminer 3y agoMy biggest project right now is training a multi-label ViT-L/16 model for a few hundred million samples. Mostly a big experiment, so not something I want to invest serious money into. I have a 2x3090 rig as my local machine, which has been useful for early experimentation, but I'm at the stage now where my runs are at 200 million samples which would take ages on that rig. 8xA100 can do it in tens of hours, which allows me to iterate faster. An 8x 3090 or 4090 machine locally would be great, but is a huge hassle to build. Last I looked into it there really wasn't a lot of knowledge available online on how to even do it. I did find an EPYC server motherboard and such that I could theoretically use, but couldn't find a great source for a >3,200 Watt server power supply. Everything in that domain is geared towards either B2LargeCorp, or B2ServerBuilders2SmallBusinesses. I could of course buy an 8xA100 rig no problem for some ungodly amount of money, but as noted above that's not appropriate for this project. In both cases, I now have to figure out what to do with 3,200kW of heat output in my office which I'm trying to avoid turning into sauna. Or co-lo it for more money out of pocket. So renting off the cloud has worked well enough at this scale.
- pmoriarty 3y ago"Mostly a big experiment, so not something I want to invest serious money into." So how much did it cost?
- esquire_900 3y agoYou might be interested in http://nonint.com/ http://nonint.com/. He made a number of quite detailed blogposts on his 2 gpu machines (both 8x 3090's), including racks, power delivery etc
- sbierwagen 3y ago>heavily differentiating the H100 products by their interface this go around, which I find quite odd. Power draw. The 4090 pulls 450 watts, and is destroying connectors. H100 SXM uses 700 watts. If the PCI consortium wants NVIDIA to use their standard for top end products, they need to make a better connector.
- lyu07282 3y agoIt was a cheap power adapter design that was sold by Nvidia that caused the melting, it was nvidias own doing selling a cheap cable on a $1,600 card. Using proprietary interfaces has everything to do with monopolistic behavior and nothing with faulty standards.
- nullindividual 3y ago12VHPWR is art of the ATX 3.0 standard and not proprietary to nVidia.
- lyu07282 3y agoThe adapter caused the issue not the connector itself, nobody who used power supplies with a native connector had any issues. I was talking about the SXM interface instead of PCIe
- ydau 3y agoHey! FYI, we are hoping to roll out a fix for the PyTorch issue tomorrow (I’m one of the founders of Lambda). Also, that’s good feedback on the CPUs bottlenecking. I’ll let our HPC hardware team know about this. We are also looking into GPU direct storage to help resolve this: https://developer.nvidia.com/blog/gpudirect-storage/ https://developer.nvidia.com/blog/gpudirect-storage/
- guisalberto 3y agoAs you were CPU limited on LambdaLabs, have you tried GPU instances from Latitude.sh? https://www.latitude.sh/accelerate/pricing https://www.latitude.sh/accelerate/pricing Each H100 is paired with 32 3.25GHz CPUs