10 ms·
Are GPUs Worth It for ML?
- ummonk 4y agoWhat a clickbaity article. It’s an interesting discussion of GPU multiplexing for ML inference merged together with a sales pitch but the clickbait title made me hate the article bait and switch. This wasn’t even an example of Betteridge’s law but just completely misleading headline.
- Eridrus 4y agoIs everyone with relevant inference costs not doing this already? I am so confused how there seems to be a startup around having a work queue that does batching...
- deleted 4y ago[deleted]
- Kukumber 4y agoAn interesting question, shows how insanely overpriced GPUs still are, specially in the cloud environment
- pqn 4y agoDisclaimer: I work at Exafunction I empathize a bit with the cloud providers as they have to upgrade their data centers every few years with new GPU instances and it's hard for them to anticipate demand. But if you can easily use every trick in the book (CPU version of the model, autoscaling to zero, model compilation, keeping inference in your own VPC, using spot instances, etc.) then it's usually still worth it.
- lowdose 4y agoNot to mention AWS has had a GPU cloud offering monopoly because Google Cloud and Microsoft Azure were publicly available until 2019.
- fomine3 4y agoGCP still provides NVIDIA K80. I wonder is it still worth to hold.
- varunkmohan 4y agoI think you'd probably always want to go with T4's since they are the same price unless there's just no availability for them.
- NavinF 4y ago*only in the cloud environment Throw some 3090s in a rack and you’ll break even in 3 months
- mistrial9 4y agothe HPC crowd are not able to add GPUs, that I know of.. deepLearning group of algorithms do kick butt for lots of kinds of problems+data .. though I will advocate that dl is NOT the only game in town, despite what you often read here
- Frost1x 4y agoIn what context? HPC and certain code bases have been effectively leveraging heterogenous CPU GPU workloads for a variety of applications for quite awhile. I know of some doing so in at least 2009 and know plenty of prior art was already there by that point, it's just a specific time I happen to remember.
- mistrial9 4y agook - the academic study in front of me dated 2020 says "no" but it is non-US researchers, public science. I have no reason to believe one way or the other, but I literally read this today. reading again - it seems this paper calls HPC with GPUs a slightly different name "GPGPU" and lists the research activity separately.. so I didn't see it as HPC; basically what I wrote is not accurate. got it
- 37ef_ced3 4y agoFor small-scale transformer CPU inference you can use, e.g., Fabrice Bellard's https://bellard.org/libnc/ https://bellard.org/libnc/ Similarly, for small-scale convolutional CPU inference, where you only need to do maybe 20 ResNet-50 (batch size 1) per second per CPU (cloud CPUs cost $0.015 per hour) you can use inference engines designed for this purpose, e.g., https://NN-512.com https://NN-512.com You can expect about 2x the performance of TensorFlow or PyTorch.
- tombert 4y agoIs there a thing that Fabrice Bellard hasn't built? I had no idea that he was interested in something like machine learning, but I guess I shouldn't have been surprised because he has built every tool that I use.
- mistrial9 4y agohttps://en.wikipedia.org/wiki/Fabrice_Bellard https://en.wikipedia.org/wiki/Fabrice_Bellard
- deleted 4y ago[deleted]
- nl 4y agoIf you are in the "data compression ~= intelligence" camp then Fabrice Bellard is currently leading the race to AI too. http://prize.hutter1.net/ http://prize.hutter1.net/ https://bellard.org/nncp/ https://bellard.org/nncp/ http://www.mattmahoney.net/dc/text.html http://www.mattmahoney.net/dc/text.html
- deleted 4y ago[deleted]
- PeterisP 4y agoFor some reason they focus on the inference, which is the computationally cheap part. If you're working on ML (as opposed to deploying someone else's ML) then almost all of your workload is training, not inference.
- varunkmohan 4y agoAgreed that there are workloads where inference is not expensive, but it's really workload dependent. For applications that run inference over large amounts of data in the computer vision space, inference ends up being a dominant portion of the spend.
- PeterisP 4y agoThe way I see it, generally every new data point (on which the production model inference gets run once) becomes part of the data set which then gets used in training every next model, processing the same data point many more times in training, thus training unavoidably taking more effort than inference. Perhaps I'm a bit biased towards all kinds of self-supervised or human-in-the-loop or semi-supervised models, but the notion of discarding large amounts of good domain-specific data that get processed only for inference and not used for training afterward feels a bit foreign to me, because you usually can extract an advantage from it. But perhaps that's the difference between data-starved domains and overwhelming-data domains?
- pdpi 4y agoThere's one piece of the puzzle you're missing: field-deployed devices. If I play chess on my computer, the games I play locally won't hit the Stockfish models. When I use the feature on my phone that allows me to copy text from a picture, it won't phone home with all the frames.
- varunkmohan 4y agoYup, exactly. It's a good point that for self-supervised workloads, the training set can become arbitrarily large. For a lot of other workloads in the vision space, most data needs to be labeled to be able to used for training.
- sabotista 4y agoIt depends a lot on your problem, of course. Game-playing (e.g. AlphaGo) is computationally hard but the rules are immutable, target functions (e.g., heuristics) don’t change much, and you can generate arbitrarily sized clean data sets (play more games). On these problems, ML-scaling approaches work very well. For business problems where the value of data decays rapidly, though, you probably don’t need the power of a deer or complex neural net with millions of parameters, and expensive specialty hardware probably isn’t worth it.
- scosman 4y agoWe did a big analysis of this a few years back. We ended up using a big spot-instance cluster of CPU machines for our inference cluster. Much more consistently available than spot GPU, at greater scale, and at better price per inference (at least at the time). Scaled well to many billion inferences. Of course, compare cost per inference on your models to make sure logic applies. Article on how it worked: https://www.freecodecamp.org/news/ml-armada-running-tens-of-billions-of-ml-predictions-on-a-budget-f9505c820203/ https://www.freecodecamp.org/news/ml-armada-running-tens-of-... Training was always GPUs (for speed), non-spot-instance (for reliability), and cloud based (for infinite parallelism). Training work tended to be chunky, never made sense to build servers in house that would be idle some of the time, and queued at other times.
- varunkmohan 4y agoDisclaimer: I'm the Cofounder / CEO at Exafunction That's a great point. We'll be addressing this in an upcoming post as well. We've served workloads that run entirely on spot GPUs where it makes sense since a small number of spot GPUs can make up for a large amount of spot CPU capacity. The best of all worlds is if you can manage both spot and on-demand instances (with a preference towards spot instances). Also, for latency sensitive workloads, running on spot instances or CPUs sometimes is not an option. I could definitely see cases where it makes sense to run on spot CPUs though.
- fortysixdegrees 4y agoDisclaimer != Disclosure Probably one of HNs most common mistakes in comments
- adgjlsfhk1 4y agoperhaps, but I think disclaimer in this context it's just an abbreviation since the disclosure carries with it the implicit disclaimer of "so the things I'm saying are subconsciously influenced by the fact that they potentially could make me money"
- fragmede 4y ago
- jacquesm 4y agoThis is an ad.
- mpaepper 4y agoThis also very much depends on the inference use case / context. For example, I work in deep learning on digital pathology where images can be up to 100000x100000pixels in size and inference needs GPUs as it's just way too slow otherwise.
- PeterStuer 4y ago" It feels wasteful to have an expensive GPU sitting idle while we are executing the CPU portions of the ML workflow" What is expensive? Those 3090ti's are looking very tasteful at current prices.
- fancyfredbot 4y agoThere are some pretty elegant solutions out there for the problem of having the right ratio of CPU to GPU. One of the nicer ones is rCUDA. https://scholar.google.com/citations?view_op=view_citation&hl=es&user=4XgrRlMAAAAJ&citation_for_view=4XgrRlMAAAAJ:zYLM7Y9cAGgC https://scholar.google.com/citations?view_op=view_citation&h...
- varunkmohan 4y agorCUDA is super cool! One of the issues though is for a lot of the common model frameworks are not supported and a new release has not come out a while.
- fancyfredbot 4y agoFair point. It's not obvious from the website which model frameworks does exafunction supports, or when the last exafunction release was.
- varunkmohan 4y agoYeah, we should have a public release very soon for people to deploy internally. We will have support for all the commonly-used frameworks and different versions.
- fancyfredbot 4y agoSounds awesome, look forward to it.
- andrewmutz 4y agoAt training time they sure are. The only thing more expensive than fancy GPUs are the ML engineers whose productivity that are improving.
- rfrey 4y agoNot related to the article, but how would one begin to become smart on optimizing GPU workloads? I've been charged with deploying an application that is a mixture of heuristic search and inference, that has been exclusively single-user to this point. I'm sure every little thing I've discovered (e.g. measuring cpu/gpu workloads, trying to multiplex access to the gpu, etc) was probably covered in somebody's grad school notes 12 years ago, but I haven't found a source of info on the topic.
- pqn 4y agoLet's just take the topic of measuring GPU usage. This alone is quite tricky -- tools like nvidia-smi will show full GPU utilization even if not all SMs are running. And also the workload may change behavior over time, if for instance inputs to transformers got longer over time. And then it gets even more complicated to measure when considering optimizations like dynamic batching. I think if you peek into some ML Ops communities you can get a flavor of these nuances, but not sure if there are good exhaustive guides around right now.
- einpoklum 4y ago> And CPUs are so much cheaper Doesn't look like it. Consumer: AMD ThreadRipper 3970X: ~3000 USD on NewEgg https://www.newegg.com/amd-ryzen-threadripper-2990wx/p/N82E16819113618?Description=AMD%20Ryzen&cm_re=AMD_Ryzen-_-19-113-618-_-Product&quicklink=true https://www.newegg.com/amd-ryzen-threadripper-2990wx/p/N82E1... NVIDIA RTX 3080 Ti Founders' Edition: ~2000 USD https://www.newegg.com/nvidia-900-1g133-2518-000/p/1FT-0004-006T6?Description=Geforce%20RTX%203080%20Ti%20Founders%20edition&cm_re=Geforce_RTX%203080%20Ti%20Founders%20edition-_-1FT-0004-006T6-_-Product&quicklink=true https://www.newegg.com/nvidia-900-1g133-2518-000/p/1FT-0004-... For servers, a comparison is even more complicated and it wouldn't be fair to just give two numbers, but I still don't think GPUs are more expensive. ... besides, none of that may matter if yours is a power budget.
- fomine3 4y agoConsumer GPUs are very cheap but prohibited to use it on datacenter.
- einpoklum 4y agoIn a datacenter you need to compare Xeon's and Epyc's with Telsa's.
- fomine3 4y agoTechnically no one prevents using Core/Ryzen series on datacenter.
- atq2119 4y agoThat only applies to Nvidia GPUs.
- fomine3 4y agoAh yes but recent consumer RADEONs are not suitable for computing task (ROCm still experimental?), while Geforce is always fine for FP32 or below.
- synergy20 4y agoI think TPU is the way to go for ML, be it training or inference. We're using GPU(some contains a TPU block inside) due to 'historical reasons'. With vector unit(x86 AVX, ARM SVE, RISC-V RVV) that is part of the host cpu, either put a TPU on a separate die of the chiplet, or just put it into a PCIe card will do the heavy lift ML job fine. It shall be much cheaper than the GPU model for ML nowadays, unless you are both a PC game player and a ML engineer.
- why_only_15 4y agoIt's true that we were initially using GPUs mostly for historical reasons, but over the last several years modern GPUs have been optimized for ML as much as anything else. If you read NVidia's marketing documents, they talk constantly about ML. The A100 is about as good, if not better, than the TPUv4 in terms of raw performance on ML workloads. The A100 can do 312 bf16 TFLOPs and costs $0.88/hr on Google Cloud [0] whereas the TPUv4 can do 275 bf16 TFLOPs and costs $0.97/hr on Google Cloud [1] [2]. The A100 is also generally speaking easier to program: it's supported by more frameworks and can perform more operations. The TPUv4 is in my understanding still worth it if you like JAX and/or you're doing lots of networking though. WRT putting a TPU on a separate die -- this has been done for several years in the mobile space: Apple Neural Engine for iPhones, TPU (not same as server TPU) on Pixel, SNPE on Qualcomm, etc. [0] https://cloud.google.com/compute/gpus-pricing https://cloud.google.com/compute/gpus-pricing [1] https://cloud.google.com/tpu/pricing#v4-pricing https://cloud.google.com/tpu/pricing#v4-pricing [2] this is somewhat unfair, because the GPU pricing number is for just the GPU and not the host it runs on, whereas the TPU pricing number (for TPU VMs) includes the host it runs on. If you include the price GCP charges for the host, preemptible A100s are about $1.20/hr. Why does Google make GPUs look cheaper than TPUs when they're not? Your guess is as good as mine.
- synergy20 4y agoMaybe Google is favoring TPUv4 over whatever GPU runs on its platform? With Hopper 100 on the way, I wonder when TPUv5 will come out. I also wonder how Intel's Gaudi2 vs Ponte Vecchio will work together, looks like duplicate efforts for me. AMD has its MI300 on the way, but it seems still far behind Nvidia|TPU|Intel at this point.
- triknomeister 4y agoI thought this post would be about how ASICs are probably a better bet.
- rvz 4y agoNot only the end result of these deep learning models can be tricked over a single pixel or get confused by malicious input and becomes useless, Deep Learning training, retraining, fine tuning on GPUs, TPUs, all running in a data center contribute significantly to burning up the planet and driving up costs which the models are just used for nothing but surveillance on our own data. If it doesn't work it has to be retrained on new data again and there are no efficient alternatives to this energy waste other than use more GPUs, TPUs, etc emitting more CO2 after years of Deep Learning existing. A complete waste of resources and energy. Therefore it is not worth it at all.
- visarga 4y agoWhy so negative? It's a waste only if the value provided is less than the cost. You can't decide that with only the cost. As humans we have our own adversarial examples, we get tired, we get sloppy, we might be even more biased than a calibrated model and always much more expensive.
- rvz 4y ago> Why so negative? It's a waste only if the value provided is less than the cost. It is entirely true and it just takes an invalid input to trick them and it messes up easily and even worse when there are always biases involved. Thus the value is nullified. And once that model breaks and doesn't work, what is the solution? More retraining on new data? Even with that like I said there are ZERO efficient alternatives, which the cost outweighs the benefits. Therefore, it is not even worth it.
- jimmygrapes 4y agoPerhaps it's been mentioned before but I do find it curious how often crypto mining was lambasted for contributing to climate change get I haven't seen anybody bat an eye at a fairly similar amount of compute power used for ML applications. Makes me wonder.
- atq2119 4y agoThe two are quite different when you look at the cost/benefit ratio.
- deleted 4y ago[deleted]