10 ms·
Cloud GPU Resources and Pricing
- rlupi 3y agoOut of curiosity. Do mostly use/want one GPU or the full server with all GPUs (8x A100 80GB, or 16x A100 40GB but I think only Google Cloud has those)? or a mix?
- nihit-desai 3y agoA comprehensive list of GPU options and pricing from cloud vendors. Very useful if you're looking to train or deploy large machine learning/deep learning models.
- tedivm 3y agoI love the full stack deep learning crew, and took their course several years ago in Berkeley. I highly recommend it. One thing that always blows my mind is how much it is just not worth it to train LLMs in the cloud if you're a startup (and probably even less so for really large companies). Compared to 36 month reserved pricing the break even point was 8 months if you bought the hardware and rented out some racks at a colo, and that includes the on hands support. Having the dedicated hardware also meant that researchers were willing to experiment more when we weren't doing a planned training job, as it wouldn't pull from our budget. We spent a sizable chuck of our raise on that cluster but it was worth every penny. I will say that I would not put customer facing inference on prem at this point- the resiliency of the cloud normally offsets the pricing, and most inference can be done with cheaper hardware than training. For training though you can get away with a weaker SLA, and the cloud is always there if you really need to burst beyond what you've purchased.
- ericpauley 3y ago> Compared to 36 month reserved pricing the break even point was 8 months if you bought the hardware and rented out some racks at a colo, and that includes the on hands support. Having the dedicated hardware also meant that researchers were willing to experiment more when we weren't doing a planned training job, as it wouldn't pull from our budget. That’s having it both ways, of course. You can’t both recoup the hardware cost in 8 months and have “free” downtime. Under this pricing you need at least 25%ish duty cycle to break even (in 3 years) so probably still favoring buying, but for some people that might not add up. Pricing also varies drastically between providers, so this may depend on choice there.
- rfoo 3y agoThe thing is, this is compared to 36 months reserved instance, which also assumes 100% usage. For true on-demand you pay much more.
- moonchrome 3y agoReserved pricing is not on-demand.
- bradleyjg 3y agoIt’s the first widespread thing where build your own makes sense in a while. Most prior workflows either needed high SLAs, were bursty, or just didn’t add up to much.
- dkobran 3y agoFWIW, Paperspace has a similar GPU comparison guide located here https://www.paperspace.com/gpu-cloud-comparison https://www.paperspace.com/gpu-cloud-comparison Disclosure: I work on Paperspace
- slavik81 3y agoI was very hopeful when I saw "AMD Support" listed on a few of those providers, but that appears to only refer only to AMD CPUs. It is, unfortunately, very difficult to find public cloud providers for AMD GPUs.
- sva_ 3y agoIt seems very honest of you to release this list, considering you're at the very bottom of it by $/hour.
- akiselev 3y agoDemand outstrips supply by such a wide margin it doesn’t matter one bit how they compare. The cheapest 8x A100 (80GB) on the list is LambdaLabs @ $12/hour on demand, and I’ve only once seen any capacity become available in three months of using it. AWS last I checked was $40/hr on demand or $25/hr with 1 year reserve, which costs more than a whole 8xA100 hyperplane from Lambda. The pricing on these things is nuts right now
- pavelstoev 3y agoVery nice. Can we chat briefly ?
- KeplerBoy 3y agowhy does no one rent out AMD GPUs? I know those cards are second class citizens in the world of deep learning, but they have had (experimental) pytorch support for a while now, where are the offerings?
- detaro 3y agoWhy would anyone offer a strictly worse product unless it were lots cheaper to offer, which it isn't? (Even for non-AI use cases, I don't think AMD has much that's more attractive in servers?)
- mkaic 3y ago> I know those cards are second class citizens in the world of deep learning, It's worse than that. AMD cards aren't second class citizens, they're not even on the same playing field. ROCm can't compete with CUDA and its ecosystem at all, the most popular deep learning frameworks are only experimentally supported, and Nvidia ships more dedicated tensor processing cores for AI acceleration on their cards. Nvidia has a near monopoly in AI not because they're particularly amazing, but because it seems like AMD is just uninterested in competing.
- arvinsim 3y agoI don't think that AMD is uninterested in competing. It is just that the mindshare is swallowed up by Nvidia that it is really difficult to use something else even if you want to.
- kouteiheika 3y ago> I don't think that AMD is uninterested in competing. I also think they're not interested. Either that or just simply incompetent. For example, just look at this issue and see the huge mess: https://github.com/RadeonOpenCompute/ROCm/issues/1714 https://github.com/RadeonOpenCompute/ROCm/issues/1714 With NVidia I can just buy any random GPU and expect it to work for everything I throw at it (at long as it has enough VRAM). With AMD it's a roulette, and only a handful of very expensive server/workstation GPUs (8 in total if I'm counting it right) are actually officially supported. It's a joke. They need to better support their own products, and they need to officially support all of their consumer GPUs to expand their mindshare. They're not doing that. From what I can see they only seem to be interested in the traditional HPC space.
- stuckkeys 3y ago"Those god damn AWS charges" -Silicon Valley. Might as well build your own GPU farm. Some of these cards, used you can probably get for 6K (guestimating).
- x-complexity 3y agoThat would imply that the current AI cycle would be able to persist at its current levels of frothiness indefinitely: In the in-between lull periods, these GPU farms would be seen as something to sell off. This doesn't even take into account the eventual depreciation of the GPUs in question, as better GPUs/accelerators come into the market. Most companies have an AWS account that they can throw on more money at for 'AI research & implementation'. With such an account existing in the first place, along with said price depreciations, the company in question would have to be certain that they'll use said GPUs all the time to make up for the upfront costs they'll be putting up with.
- stuckkeys 3y agoSure bud. Sure. To each their own. When these GPU enter the consumer market is when the AWS cost becomes irrelevant. =)
- mirekrusin 3y agoYour own hardware can be rented out if you're not using it through vast.ai for example. When new hardware comes out, you can sell old one to recover some of the cost.
- pavelstoev 3y agoOften you are better off running certain workloads on lesser GPUs. But this requires certain tricky compiler-level optimizations. For example, can run certain LLM inference with comparable latency on cheaper A40s vs running on A100s. Could also run on 3090s (sometimes even faster). This helps with operating costs but may also resolve availability constraints.
- sorcer 3y ago
- andrewstuart 3y agoGo buy a GPU from the local computer store. Consumers GPUs are much more available, much cheaper and much faster.
- throwaway3672 3y agoWhy just not rent a 3090 from vast ai? This is literally 0.2$ per hour.
- deleted 3y ago[deleted]
- vhcr 3y agoYou can purchase a 3090 for $1000, assuming you're going to use it 24/7, at 450W you would use about 1000kWh, at $0.10 / kWh, it would pay itself in about 3 months.
- arvinsim 3y agoThe value proposition is very different between one who is just starting to learn AI/ML and one who already knows what to train.
- dikei 3y agoThat's not a typical use case during development though. People need fast feedback loops: they'd rather rent multiple GPUs and have the result the next morning, than waiting for days with no guarantee of success. So unless you have stable tasks that need to run continuously, or have enough users to keeps your GPU clusters busy, your GPUs usage would be quite bursty: some period of high activity then a lot of time idling.
- oceanplexian 3y agoI'd argue the opposite. Having to spin up and down instances for development is a huge PITA, the tooling sucks, and the instance might not even be available the next time you need them. It also stresses me out personally, because I'm worrying about getting productive use out of every minute. Whereas my little GPU cluster (Despite a big upfront cost) costs nothing but electricity and runs 24/7.
- hislaziness 3y agoI do not see any of the AI processors like the Google TPU. Would they be cheaper?
- brucethemoose2 3y agoTPU and other accelerator performance varies by application... And even in a hugely popular config (like finetuning LLaMA with JAX) its hard to find a good benchmark. But generally speaking Google charges a pretty penny for TPUs Accelerators outside TPUs are exotic. Off the top of my head... Cerebras only offers their WS2 as a 1st party "pay for a specific training job" kinda thing. Intel Gaudi 2 is supposedly good but is mysterious to me, and Ponte Vecchio has barely started shipping. Graphcore and Tenstorrent chips in the wild seem kinda long in the tooth for big training jobs. The AMD MI300 is not shipping, and the older AMD Instincts are difficut to find in cloud services (maybe because they got eaten up for HPC?) Lots of other promising accelerators (with my personal favorite being the Centaur x86 "accelerated" CPUs, perfect for dirt cheap LLM inference/LORA training) died on the vine because of the CUDA moat, and I think more will share the same fate.
- pavelstoev 3y agoMay also suggest the suite of open source DeepView tools and which are part of PyPi. Profile and predict your specific model training performance on a variety of GPUs. I wrote a linked in post with usage GIF here: https://www.linkedin.com/posts/activity-7057419660312371200-hdb5?utm_source=share&utm_medium=member_desktop https://www.linkedin.com/posts/activity-7057419660312371200-... And links to PyPi: => https://pypi.org/project/deepview-profile/ https://pypi.org/project/deepview-profile/ => https://pypi.org/project/deepview-predict/ https://pypi.org/project/deepview-predict/ And you can actually do it in browser for several foundational models (more to come): => https://centml.ai/calculator/ https://centml.ai/calculator/ Note: I have personal interests in this startup.
- freediver 3y agoCreated this a while ago for the same purpose https://cloudoptimizer.io https://cloudoptimizer.io
- justherefornews 3y agoNice man, ty!
- aeturnum 3y agoThis isn't my area of expertise so I can't get too deep but this is an extremely well organized and straightforward resource.
- IceHegel 3y agoDo we have numbers from H100
- sorcer 3y agoThe only H100s available to get is the "little brother" PCIe version with HBM2e memory instead of HBM3. Very very few clouds, or people, in general, have H100 deployed at any scale. You will see when the MLPerf benchmark results come out.
- fpgaminer 3y agoMy experience with many of these services renting mostly A100s: LambdaLabs: For on-demand instances, they are the cheapest available option. Their offering is straightforward, and I've never had a problem. The downside is that their instance availability is spotty. It seems like things have gotten a little better in the last month, and 8x machines are available more often than not, but single A100s were rarely available for most of this year. Another downside is lack of persistent storage, meaning you have to transfer your data every time you start a new instance. They have some persistent storage in beta, but it's effectively useless since it's only in one region and there's no instances in that region that I've seen. Jarvis: Didn't work for me when I tried them a couple months ago. The instances would never finish booting. It's also a pre-paid system, so you have to fill up your "balance" before renting machines. But their customer service was friendly and gave me a full refund so shrug. GCP: This is my go-to so far. A100s are $1.1/hr interruptible, and of course you get all the other Google offerings like persistent disks, S3, managed SQL, container registry, etc. Availability of interruptible instances has been consistently quite good, if a bit confusing. I've had some machines up for a week solid without interruption, while other times I can tear down a stack of machines and immediately request a new one only to be told they are out of availability. The downsides are the usual GCP downsides: poor documentation, sometimes weird glitches, and perhaps the worst billing system I've seen outside of the healthcare industry. Vast.ai: They can be a good chunk cheaper, but at the cost of privacy, security, support, and reliability. Pre-load only. For certain workloads and if you're highly cost sensitive this is a good option to consider. RunPod: Terrible performance issues. Pre-load only. Non-responsive customer support. I ended up having to get my credit card company involved. Self-hosted: As a sibling comment points out, self hosting is a great option to consider. In particular "Having the dedicated hardware also meant that researchers were willing to experiment more". I've got a couple cards in my lab that I use for experimentation, and then throw to the cloud for big runs.
- pavelstoev 3y agoplease consider CoreWeave - great experience so far.
- sorcer 3y ago
- 1024core 3y agoDoesn't anybody use TPUs from Google? Given the heterogeneous nature of GPUs, RAM, tensor cores, etc. it would be nice to have a direct comparison of, say, number of teraflops-hour per dollar, or something like that.
- crucifiction 3y agoMissing Oracle Cloud which has a massive GPU footprint - https://www.oracle.com/cloud/compute/gpu/ https://www.oracle.com/cloud/compute/gpu/
- ranguna 3y agoIt's missing quite a few vendors, but they are open to PRs :)
- dotBen 3y agoLooking to run a cloud instance of Stable Diffusion for personal experimentation. Looking at cloud mostly because I don't have a GPU or desktop hardware at home, and my Mac M1 is too slow. But also needing to contend with constant switching on/off the instance several times a week to use it. Wondering which vendors other HN'ers are using to achieve this?
- Shakahs 3y agoAWS spot g5 or Vultr fractional A40 with an Ansible playbook to set up Stable Diffusion, drivers, etc.
- victorbjorklund 3y agoHave you tried google colab? Runs pretty well on it
- MacsHeadroom 3y agoIt's interesting that AWS is a full 5x more expensive than the leading low cost providers, with Google close behind AWS.
- hcarlens 3y agoI built a slightly less detailed version of this, which also also lists free credits: https://cloud-gpus.com/ https://cloud-gpus.com/ Open to any feedback/suggestions! Will be adding 4090/H100 shortly.
- TradingPlaces 3y agoLambda has a very interesting benchmark page https://lambdalabs.com/gpu-benchmarks https://lambdalabs.com/gpu-benchmarks If you look through the throughput/$ metric, the V100 16GB looks like a great deal, followed by H100 80GB PCIe 5. For most benchmarks, the A100 looks worse in comparison