11 ms·
My Experience and Advice for Using GPUs in Deep Learning: Which GPU to get
- ageitgey 8y agoThis is a great article and I highly respect his opinions. However, since you are probably eagerly reading this to see how fast the new RTX cards are, so you should know upfront that the numbers he has so far are just estimates based on specs: > Note that the numbers for the RTX 2080 and RTX 2080 Ti should be taken with a grain of salt since no hard performance numbers existed. I estimated performance according to a roofline model of matrix multiplication and convolution under this hardware together with Tensor Core benchmarks from the V100 and Titan V.
- nolok 8y agoA great way to turn a listing you can trust enough to use as one of your comparison basis, into a listing made up of imaginary marketing numbers. I guess the click baiting is needed / the best option, but I hate that's it's what most web resources are like now.
- r1nkgrl 8y agoBut the article isn't hiding the fact that the numbers are estimates. People are curious how the new cards will stack up, and this article provides the best evaluation of that given the information they have available.
- alkonaut 8y agoNo one minds comparing some products as guesses/estimates/extrapolations with some products as being real performance figures. So long as it's clear which products have which type of figure.
- p1necone 8y agoThe clock rates, number of CUDA cores, memory size/type etc in the new cards aren't really "imaginary marketing numbers". NVidia could have changed their hardware so they could put bigger numbers on paper without corresponding real world performance gains, but that's a big assumption for you to seemingly take as fact.
- shaklee3 8y agoI'd guess that the performance could be slightly better than the 1080 scaled by cores/MHz/FLOPS. The reason being that the memory bandwidth is higher on the 2080, and that's hard to model unless the person knows exactly how efficient the kernel is and if it's memory bound.
- steve_musk 8y agoPlus the architecture improvements. Do we know how many cores per SM? They’ve decoupled int and FP execution units which could give larger increases for certain kernels (although FP heavy deep learning kernels aren’t likely to benefit as much, they will still get address calculation benefits).
- shaklee3 8y agoI hadn't seen the whitepaper on Turing yet. Where did you see they decoupled them?
- steve_musk 8y agoThe keynote
- pirocks 8y agoSeems down for me: https://web.archive.org/web/20180821173206/http://timdettmers.com/2018/08/21/which-gpu-for-deep-learning/ https://web.archive.org/web/20180821173206/http://timdettmer...
- fermienrico 8y agoThe cost/performance plot - shouldn't it be "Lower is better"? It says "Higher is better". Lower value would indicate lower cost per unit level of performance. It should be "Lower is better" or the plot needs to say "Performance/Cost". Am I missing something?
- wmf 8y agoYou're missing the principle of charity.
- fermienrico 8y agohuh?
- KSS42 8y agoDo you mean "Figure 3: Normalized performance/cost numbers"? Its performance/cost and not cost/performance. Or maybe the author fixed a typo?
- songgao 8y agoI think it used to be cost/performance and was later fixed. GP left comment before the fix.
- timdettmers 8y agoThanks for your feedback! Someone mentioned this on twitter as well and I thought it was a good point so I implemented that change.
- fermienrico 8y agoThanks for being receptive. I wouldn’t call it a “good point” if it was a mistake that was corrected.
- syntaxing 8y agoHacker news hug of death? Anyone here have any experience using AMD cards with something like PlaidML? I have a 1050Ti SSC but I'm starting to feel the limitation as my complexity grows. But getting a 1080 is a bit out of my budget right now. I'm tempted to get the new Vega 56 released recently.
- steve_musk 8y agoYou could wait and see how pascal prices fall after Turing comes out.
- dostres 8y agoAn open question for me is the performance of two 2080tis using NVLink as one virtual GPU. I imagine it’ll be close to linear, but I’ll be interested to know for sure.
- shaklee3 8y agoIt won't be linear for memory-bound applications. The v100 was able to make it close to linear with large enough transfer sizes, but it has 50% more memory bandwidth than these.
- scottlegrand2 8y agoThe biggest advance here is that Nvidia has produced a consumer card that has all the high-end deep-learning features. This was missing in both the Pascal and Volta Generations even though in Pascal fp32 was full power. I think the TPU scared them and that's a good thing.
- lern_too_spel 8y agoThe "I have almost no money" recommendation should include Colab. https://medium.com/deep-learning-turkey/google-colab-free-gpu-tutorial-e113627b9f5d https://medium.com/deep-learning-turkey/google-colab-free-gp... Somebody who has almost no money isn't going to be able to equip a desktop with a GTX 1050 Ti ($175), fast disk ($50), and RAM ($50) on an entry level cpu/motherboard/power supply/case/monitor/peripherals ($300) and pay for the electricity used during training. Colab can be accessed from a free public computer or a cheap Chromebook ($200).
- deleted 8y ago[deleted]
- andy_ppp 8y agoWhat are the rules about datasets I upload to this free service? Do Google now own them?
- ColanR 8y agoMy guess it's the standard caveat: if you don't pay for the service, you and your stuff is the commodity.
- lern_too_spel 8y agoI would imagine no more than it owns the files you upload to Google Drive. The disk on the Colab instance is ephemeral, so you will need external storage for your dataset anyway.
- gaius 8y agoA cheap (but not free) option is leadergpu.com - no affiliation, they just seem like super nice people and have per-minute billing. They are Dutch.
- abcdefgh214 8y agoIf you have the programming skills necessary to develop deep learning applications, it should be assumed that you can also easily get a well-paying job so this isn't really even relevant.
- sabalaba 8y agoThe 2080Ti numbers are likely going to be a lot lower than that. We’ve benched the 1080Ti vs the Titan V and the Titan V is nowhere near 2x faster at training than the 1080Ti as suggested in that graph. We observed a 30% to 40% speedup during our benchmarking: https://deeptalk.lambdalabs.com/t/benchmarking-the-titan-v-volta-gpu-with-tensorflow/108 https://deeptalk.lambdalabs.com/t/benchmarking-the-titan-v-v... This is consistent with the 32% increase in FP32 flops from 11.3TFlops for the 1080Ti to 15TFlops for the Titan V. Additional speedups can be explained by the increase in memory bandwidth for HBM2 and the mixed precision fused multiply adds provided by the TensorCores. Thus, given the quoted 13Tflop numbers for the 2080Ti, I would expect the 2080Ti to present something more like a 15-20% speedup over the 1080Ti. So 2080Ti is less bang for your buck. But benchmarking is the only way to tell what’s better on a FLOPS/$ basis.
- timdettmers 8y agoYour data are inconsistent with the benchmarks that I mention in the blog post: https://github.com/u39kun/deep-learning-benchmark https://github.com/u39kun/deep-learning-benchmark You also do not benchmark LSTMs: https://www.xcelerit.com/computing-benchmarks/insights/benchmarks-deep-learning-nvidia-p100-vs-v100-gpu/ https://www.xcelerit.com/computing-benchmarks/insights/bench... If you put both of those benchmarks together my conclusion is quite reasonable. But I see that you could also come to your conclusion with your benchmarks. It is just a question which benchmarks are less biased and that is too difficult to evaluate. I guess we have to wait for real data, but thanks for putting your data out there to get a discussion going.
- KayL 8y agoGood article, but as a new learner, I'm interested in (your experiences on) how much time taken for the common task to train a model? 1min vs 2mins, probably I will get a cheaper GPU but if there's 5h vs 10h or 1 day vs 2 days, I'd save more money for one with good performance