9 ms·
Are there articles detailing the tradeoffs between an A100 system and a consumer system in terms of perf/$? E.g, if you're a startup trying to be scrappy and yo
by sxp 4y ago
Are there articles detailing the tradeoffs between an A100 system and a consumer system in terms of perf/$? E.g, if you're a startup trying to be scrappy and your choice is between a $100k+ A100 system and an off-the-shelf RTX 40x0 networked cluster, what would be the tradeoff?. For embarrassingly parallel use cases like ray-tracing or a compute backend that services many clients and doesn't need FP64, the consumer system would win. For other uses cases that have a hard requirement for the NVLinked access to 80GB of RAM per card, the A100 system would win.
Has anyone done benchmarks to show what the tradeoff calculations are for various use cases?
- 0x008 4y agoNot for whole systems but the single GPUs (A100/A40) are significantly better in performance per watt (almost by a factor of 2) than RTX cards for Machine Learning workloads. You can watch this LTT video to get an idea: https://youtu.be/zBAxiQi2nPc https://youtu.be/zBAxiQi2nPc Probably if you run them full time in a enterprise environment the power consumption will be the bulk of your cost?
- synergy20 4y agoI was trying to build a ML training machine myself these days. I don't really need the RTX FP64 and its 4K graphic rendering with gaming stuff, all I need is ML training(FP16 mostly with lots of CUDA cores). There is no such thing, you either use multiple RTX GPU cards and 'bend' them for ML training, or you buy A100 40GB at about 125K or A100 80GB at about 250K, which is far beyond my budget. Google's TPU could be used to assist ML training, but then it does not have the CUDA ecosystem for software, plus, it kept its new TPU for inhouse use only as far as I can tell. AWS ML EC2 are very expensive to me which is why I want to build one for mid-sized ML trainings.
- Patrick-STH 4y agoThere are a few big ones: - The CUDA license does not allow you to use GeForce in the data center. In the US it has become less popular, but if you look at our Inspur AIStation piece, that was a cluster located in China with GeForce cards. So it still happens, but less so. - The memory capacity is another big challenge. Newer models have 80GB which dwarfs the 24GB on a 4090. We just got the RTX 6000 Ada in, so that is an option for more memory. - For higher-end training, one of the big challenges is interconnect, so having NVLink and Infiniband or 100GbE+/ Infiniband NICs is important. The HGX A100 platform is designed for that with its NVSwitch and PCIe switch topology. With all of that said, you are 100% right that many startups have used consumer cards for years. For example, Andrej Karpathy talked about how our DeepLearning11 build (8x 1080 Ti's) had a ~3 month payback period versus AWS https://twitter.com/karpathy/status/924340245478256640 https://twitter.com/karpathy/status/924340245478256640
- p1esk 4y agoAndrej Karpathy talked about how our DeepLearning11 build (8x 1080 Ti's) had a ~3 month payback period versus AWS In 2017. Currently you can rent 8xA100 server for $8.8/hr: https://lambdalabs.com/service/gpu-cloud https://lambdalabs.com/service/gpu-cloud At this price the payback stretches to about 3 years (taking into account average energy costs in US, and assuming 24/7 operation for the whole 3 years).
- TOMDM 4y agoA bit of extra context, that's $8.8/hr for the 40GB A100's The 80GB A100's will run you $12.0/hr for 8 of them.
- fancyfredbot 4y agoThat's 24k over 3 months which would buy and power 8 4090 GPUs by my reckoning. Of course those 8 4090s wouldn't have enough memory to run chatGPT so maybe AWS is good value after all.
- TOMDM 4y agoThis is Lambda Labs' pricing, which is significantly cheaper than AWS.
- monkmartinez 4y agoTim Dettmers [0] has been updating the linked article for a while now. It is a great primer -> deep dive into the world of GPU's for ML/DL tasks. I don't need 4x RTX 4090's but I really want them. As it stands now, I use an old M40 that I modified for water cooling. The only tricky part was configuring Win10 to run the card in WDDM mode as opposed to TCC mode without onboard graphics. To be honest, I am not entirely sure how it works as the pass-through is an older quadro card. How do I justify the purchase of DGX-1 as a hobby programmer? Call it my mid-life crisis purchase? I mean, its not that much more than a mid-range 2023 Corvette Stingray. [0] https://timdettmers.com/2023/01/30/which-gpu-for-deep-learning/ https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni...
- KaiserPro 4y agoUnless you are planning to have a full system on all the time, renting a loaded 8xA100 machine is almost certainly the way to go for a startup. As others have pointed out, vRAM is the killer here. this big of a machine is mostly for vRAM coherence, which is a specific usecase. For almost every other type of job lots of single GPUs with shared disk storage will do (particle sim, rendering, oil and gas, etc.)
- paulpan 4y agoHistorically Nvidia's professional-grade and consumer products were artificially segmented. The only difference was in their amount of VRAM, with recent example is the Pascal-generation GeForce 1080 vs. the P100 that someone else has compared in detail: https://medium.com/@alexbaldo/a-comparison-between-nvidias-geforce-gtx-1080-and-tesla-p100-for-deep-learning-81a918d5b2c7 https://medium.com/@alexbaldo/a-comparison-between-nvidias-g... But in recent years Nvidia has added harder gating in between their products, the biggest being SLI or NVLink support reduced starting with Pascal and now entirely dropped in the latest Ada Lovelace generation. Effectively forcing you to buy the professional chips. A different angle - how long before Google, Microsoft, and Amazon design and deploy their own "GPUs" instead of buying from Nvidia? They've all started with their own ARM SOCs so surely it's inevitable.