12 ms·
Which GPU(s) to Get for Deep Learning
- 32gbsd 3y agoOmg that's a long read but very informative
- arvinsim 3y agoReally a shame that the 4070ti doesn't have 16GB. But I guessed it is expected that Nvidia doesn't want to cannibalize the 4080.
- teruakohatu 3y agoEvery level below the *100 series has some sort of limitation to give incentives to upgrade one or two levels. It's hard to blame nvidia when nobody seems to be trying to compete with them on the low end of ML and DL.
- Const-me 3y agonVidia has a 20 GB GPU with the same chip as 4070Ti, the model is RTX 4000 SFF. One issue is price, it costs almost twice as much. Another one is memory bandwidth, RTX 4000 SFF only delivers 320 GB/second. That is much slower than 4070Ti (504 GB/second) and slightly faster than 4060Ti (288 GB/second). Also the clock frequencies are half of 4070Ti, so the compute performance is worse.
- XCSme 3y ago> RTX 4000 SFF Max Power Consumption - 70W. Huh?
- Const-me 3y agoThe power efficiency of the RTX 4000 is awesome, but it costs performance. The 4070 Ti runs at 2.3 GHz base clock / 2.6 GHz boost, the RTX 4000 SFF only runs at 1.3 GHz base / 1.6 GHz boost. For this reason, despite the chip is the same, the compute performance of the RTX 4000 is not particularly great. 4070 Ti delivers up to 35.48 TFlops at base clock, RTX 4000 only 19 TFlops.
- XCSme 3y agoHalf the TFflops but 4 times less energy consumption (4070ti is TDP MAX 285W)
- politelemon 3y agoSo Nvidia is going to pretty much corner the market for a long time? This bit I expected but was still sad to read. Surely we would benefit from competition. It would probably take a lot of investment from AMD to make that happen, I imagine. > AMD GPUs are great in terms of pure silicon: Great FP16 performance, great memory bandwidth. However, their lack of Tensor Cores or the equivalent makes their deep learning performance poor compared to NVIDIA GPUs. Packed low-precision math does not cut it. Without this hardware feature, AMD GPUs will never be competitive. Edit: what about Intel arc GPU? Any hope there?
- ItsBob 3y ago> It would probably take a lot of investment from AMD to make that happen, I imagine Don't AMD deliberately gimp their consumer cards to prevent cannibalising the pro cards? I vaguely recall reading about that a while back. That being the case, they have already done the R&D but they chose to use the tech on the higher-margin kit, thus preventing hobbyists from buying AMD.
- lhl 3y agoA few years ago AMD split off their GPU architectures to CDNA (focused on data center compute) and RDNA (focused on rendering for gaming and workstations). This in itself is fine and what Nvidia was already doing, it makes sense to optimize silicon for each use case, but where AMD took a massive wrong turn is that they decided to stop supporting compute completely for their RDNA (and all legacy) cards. I'm not sure exactly what AMD expected to happen when doing that, especially when Nvidia continues to support CUDA on basically every GPU they've ever made: https://developer.nvidia.com/cuda-gpus#compute https://developer.nvidia.com/cuda-gpus#compute (looks like back to a GeForce 9400 GT, released in 2008)
- empyrrhicist 3y agoIts like they don't care about having a pipeline of programmers ready to use their hardware, and want to ignore most of the workstation market.
- 3y ago
- ItsBob 3y agoJust as an FYI/additional data point, I bought a 3090 FE from Ebay a few months ago for £605 including delivery. I've only just started using it for Llama running locally on my computer at home and I have to say... colour me impressed. It generates the output slightly faster than reading speed so for me it works perfectly well. The 24GB of VRAM should keep it relevant for a bit too and I can always buy another and NVLink them should the need arise.
- PeterStuer 3y agoAnyone with experience running 2 linked consumer GPU's want to chime in how good this works in practice?
- Filligree 3y agoYou get a fast link between the GPUs, which should help when you’ve got a model split between them. However, that split isn’t automatic. You can’t expect to run a 40GB model on that, unless perhaps if it’s been designed for that—the way llama.cpp can split a model between the GPU and CPU, for instance. What you can do without trouble is keep more models loaded, do more things at the same time, and occasionally run the same model at double speed if it batches well.
- deaddodo 3y agoCUDA multi-GPU with NVLink is pretty well tested with shared memory space. You still want to use NCCL to optimize the allocation, but many CUDA-aware libraries (and their subsequent ML tools) are capable.
- pseg134 3y agoThis is incorrect if you are talking about 3090 or 3090ti using nvlink.
- PeterStuer 3y agoYou mean those would work like a virtual single GPU with 48GB vram?
- roenxi 3y agoEvaluating AMD GPU by their specs is not going to paint the full picture. Their drivers are a serious problem. I've managed to get ROCm mostly working on my system (ignoring all the notifications of what is officially supported, the jammy debs from the official repo seem to work on Debian testing). The range of supported setups is limited so it is quite easy to end up in a similar situation. I expect system lockups when doing any sort of model inference. From the experiences of the last few years I assume it is driver bugs. Based on their rate of improvement they probably will get there in around 2025, but their past performance has been so bad I wouldn't recommend buying a card for machine learning until they've proven that they're taking the situation seriously. Although in my opinion buy AMD anyway if you need a GPU on linux. Their open source drivers are a lot less hassle as long as you don't need BLAS.
- Const-me 3y agoROCm is not the only option, compute shaders are very reliable on all GPUs. And thanks to the Valve’s work on DXVK 2.0, modern Linux runs Windows D3D11 software just fine. Here’s an example https://github.com/Const-me/Whisper/issues/42 https://github.com/Const-me/Whisper/issues/42 BTW, a lot of BLAS in the compute shaders of that software.
- roenxi 3y agoI dunno, are they? AMD should to pay someone to put up some "how to multiply a 2x2 matrix on our GPU for the average programmer!" tutorials somewhere obvious. I saw a lot of GPU lockups before I gave up on trying and decided that it wasn't worth it. Maybe compute shaders were a thing I should have tried. To be honest, I don't know much about them because my attempts in the space were shut down pretty hard by driver bugs linked to OpenCL and ROCm. I thought it was just me for a while, but after watching George Hotz's famous meltdown trying to program on an AMD GPU I do wonder if they're underestimating the power of a few good public "how to use the damn thing" sessions. They've been pushing ROCm which would probably be great if it worked reliably.
- Const-me 3y ago
- nl 3y agoYou can tell how NVIDIA dominants the market by the fact their price/performance "curve" is almost a straight line. In a competitive market that line has distortions where one player trts to undercut the other. There are no bargains because there is almost no competitive pressure and so there is barely any distortion in that line.
- MrBuddyCasino 3y agoI suppose this is one of the reasons (besides AMD dropping the ball) they aren't even trying to be competitive in the gaming market - they can sell the same mm2 silicon area for much more to AI startups: "There's a full blown run on GPU compute on a level I think people do not fully comprehend right now. Holy cow. I've talked to a lot of vendors in the last 7 days. It's crazy out there y'all. NVIDIA allegedly has sold out its whole supply through the year. So at this point, everyone is just maximizing their LTVs and NVIDIA is choosing who gets what as it fulfills the order queue." [0] [0] https://twitter.com/Suhail/status/1683642991490269185 https://twitter.com/Suhail/status/1683642991490269185
- andrewstuart 3y agoSuhail lacks wisdom. Stop obsessing about cloud GPUs. Go buy retail GPUs. Make them work. Adapt your software. Whatever, stop whining about cloud GPUs just switch to retail. Solve the barriers/limitations. This has been the essences of computing for 60 years. Stop being intimidated by Nvidia. Get your job done in what is available. Buy AMD, buy Intel. Work out how to make your GPU thing work on them. Stop wringing your hands about how there’s no cloud GPUs when there’s a ton of cheap retail GPUs. INNOVATE. Look beyond nvidia. Stop whining. If you’ve bet your entire business on cloud GPUs then you’re a fool. Bet on retail GPUs.
- cosentiyes 3y ago> Solve the barriers/limitations I don't think this is possible when you can't pool memory on the 40 series retail GPUs.
- 3y ago
- PeterStuer 3y agoI'm sticking with nVidia for now (currently a 3090 bought secondhand of eBay) as it is the most tested/supported by far, but it is great to see AMD making progress (finally) as some competition in this segment is desperatly needed.
- fnands 3y agoAny tips for getting one off ebay without getting screwed? I want to pull the trigger, but a bit scared.
- stephenitis 3y agoditto. Second hard graphics cards are such a wild west to me.
- PeterStuer 3y agoIt was my first purchase of of ebay, so not sure I can advice much. I just waited for a reputable seller to show up. I also limited myself to professional sales from the EU to avoid any potential import issues. I guess there is always a risk involved, but probably more buying from a first time private profile with just one object listed, than from a business that sells every day with high rep.
- savandriy 3y agoI've bought a Radeon RX 6700XT (12GB) last year, primarily for playing games. But after Stable Diffusion came out, I started to play around with it and was pleasantly surprised that the GPU could handle it! The setup is a little messy, and Linux only. For someone targeting AI, definitely pick an Nvidia card with 12+ GBs of VRAM.
- xnx 3y agoDo local GPUs make sense? For the same price, can't you got a full years worth of cloud gpu time?
- disintegore 3y agoLooking at the pricing, if you only spin those instances up when you need them, you can go a while before you break even. Otherwise it only takes a few months depending on the GPU. I would imagine that someone really serious about training (or any other CUDA workload) uses both.
- wing-_-nuts 3y agoHaving looked at the pricing of retail card vs cloud, I came to the conclusion I could probably buy enough cloud compute to complete a phd before I 'paid for' the cost of a 4090 build...
- TillE 3y agoBuying a high-end gaming GPU also lets you do, well, high-end gaming, 3D and video renders, etc. If you only care about ML stuff, sure, the calculation is different.
- Yenrabbit 3y agoCloud GPU providers are running low on capacity at the moment as people frantically suck up capacity to hop on the AI bandwagon, raising worries about availability. So having guaranteed access is maybe one motivation for local GPUs. But for me the main reason to go local is more psychological. I've mostly used cloud compute up until now but whenever I'm paying an hourly cost (even a small one) there is a pressure to 'make it worthwhile' and I feel guilty when the GPU is sitting idle. This disincentivizes playing and experimentation, whereas when you can run things locally there is almost no friction for quickly trying something out.
- synergy20 3y ago4090 is now in high end PCs, with 24GB VRAM, that's what I'm going to buy. Everyone talks about Nvidia GPUs and AMD MI250/MI300, where is Intel? Would love to have a 3rd player.
- singhrac 3y agoIntel has Habana Gaudi2, which is an A100 competitor, but you can only access it on Intel’s developer cloud, apparently.
- synergy20 3y agoyes even MI300 from AMD is data center only just like A100 and H100. I guess what Intel is missing is that it does not have a PC version GPU(ARC is far behind AMD and Nvidia GPU cards), so it can not establish its developer ecosystem and its OneAPI is a hard sell for its AI plan. Either make ARC or whatever GPU as good as Nvidia/AMD graphic cards, or at least make lots of great AI compute accelerator stick to stay in the game, or no future in the AI era for Intel, sadly.
- whywhywhywhy 3y agoConsider the 3090, same memory but was way cheaper than the 4090 when I was looking, might be a good trade off if you don't really need the 40 speed boost.
- synergy20 3y ago3090 is still under powered by quite a bit though it does have 24GB
- andy_ppp 3y agoI hear a lot about CUDA and how bad ROCm is etc. and I’ve been trying to understand what exactly CUDA is doing that is so special; isn’t the maths for neural networks mostly multiplying large arrays/tensors together? What magic is CUDA doing that is so different for other vendors to implement? Is it just lock-in, the type of operations that are available, some kind of magical performance advantage or something else that CUDA is doing?
- empyrrhicist 3y ago1. Driver stability 2. Works on more consumer grade cards 3. Ecosystem advantage (lots of software developed against an existing and well supported ecosystem) I have a laptop with a mobile 2060 and a desktop with a top-of-the-line consumer 7900XTX. As of yet, the 7900XTX isn't officially supported (and I haven't bothered to go down the obnoxious rabbit hole to figure out how to compute on it). Meanwhile, I can load up CUDA.jl on my laptop in mere minutes with absolutely no fuss. Edit: if there are any GPU gurus out there who are capable of working on AMDGPU.jl to make it work on cards like the 7900XTX out of the box and writing documentation/tutorials for it... start a Patreon. I bet you could fund some significant effort getting that up and running!
- officialchicken 3y agoAs of today, there is zero consumer card support from AMD. It is an option only if you have a PRO card. "Formal support for RDNA 3-based GPUs on Linux is planned to begin rolling out this fall, starting with the 48GB Radeon PRO W7900 and the 24GB Radeon RX 7900 XTX, with additional cards and expanded capabilities to be released over time." [0] [0] https://community.amd.com/t5/rocm/new-rocm-5-6-release-brings-enhancements-and-optimizations-for/ba-p/614745 https://community.amd.com/t5/rocm/new-rocm-5-6-release-bring...
- empyrrhicist 3y agoRight, which SUCKS. Everyone who wants to prototype on their existing gear before jumping into a big pro card purchase is stuck with Nvidia, and the availability/performance of the software stack shows it.
- graton 3y agoI almost immediately became suspicious on the accuracy of this article when they said the "Nvidia RTX 40 Ampere series". Ampere was the architecture name for the RTX 30 series. Ada Lovelace is the architecture name for the RTX 40 series.
- fnands 3y agoProbably just an accident. Tim Dettmers has been updating this post for years and it's a super valuable resource.
- reducesuffering 3y agoYou'll want lots of memory, so depends on your price point. 4090 ($1,600) > 3090 ($1300 new - $600 used) > 3060 ($300) used 3090 is ideal value. Lots of models will need the 24gb ram
- pizza 3y agoTrying to build a scalable home 4090 cluster but running into a lot of confusion... Let's say - I have a motherboard + cpu + other components and they've both got plenty of pcie lanes to spare, total this part draws 250W (incl the 25% extra wattage headroom) - start off with one RTX 4090, TDP 450W, with headroom ~600W. - I want to scale up by adding more 4090s over time, as many as my pcie lanes can support. 1. How do I add more PSUs over time? 2. Recommended initial PSU wattage? Recommended wattage for each additional pair of 4090s? 3. Recommended PSU brands and models for my use case? 4. Is it better to use PCI gen5 spec-rated PSUs? ATX 3.0? 12vhpwr cables rather than the ordinary 8-pin cables? I've also read somewhere that power cables between different brands of PSUs are *not* interchangeable?? 5. Whenever I add an additional PSU, do I need to do something special to electrically isolate the PCIe slots? 6. North American outlets are rated for ~15A * 120V. So roughly 1800W. I can just use one outlet per psu whenever it's under 1800W, right? For simplicity let's also ignore whatever load is on that particular electrical circuit. Each GPU means another 600W. Let's say I want to add another PSU for every 2 4090s. I understand that to sync the bootup of multiple PSUs you need an add2psu adapter. I understand the motherboard can provide ~75W for a pcie slot. I take it that the rest comes from the psu power cables. I've seen conflicting advice online - apparently miners use pcie x1 electrically isolated risers for additional power supplies, but also I've seen that it's fine as long as every input power cable for 1 gpu just comes from one psu, regardless of whether it's the one that powers the motherboard. Either way x1 risers is an unattractive option bc of bandwidth limitations. pls help
- steffan 3y ago> 6. North American outlets are rated for ~15A * 120V. So roughly 1800W. I can just use one outlet per psu whenever it's under 1800W, right? For simplicity let's also ignore whatever load is on that particular electrical circuit. You're going to have a bad time with this assumption; typical non-kitchen household circuits in the U.S. are 15A for the circuit. Each outlet is usually limited to 15A, but the circuit breaker serving the entire circuit is almost certainly 15A as well; one outlet at maximum load will not leave capacity for another outlet on the same circuit to be simultaneously drawing maximum amperage. Typical residential construction would have a 15A circuit for 1-2 rooms, often with a separate circuit for lighting. Some rooms, e.g. kitchens will have 20A circuits, and some houses may have been built with 20A circuits serving more outlets / rooms.
- andrewstuart 3y agoThere’s clearly demand to buy AI capable GPUs at the store at a low price. But Nvidias monopoly mean a they cripple their retail cards and push the AI stuff to data centers. If only there was many manufacturers of AI hardware and software there would be abundant cheap products at every level. AMD and Intel don’t seem to be able to compete and there’s no sign that will change. So AI is going to remain expensive and hard to get for a very long time.
- fnands 3y agoApp based on this post to help you decide what to buy: https://nanx.me/gpu/ https://nanx.me/gpu/
- cosmojg 3y agoTL;DR, your best option right now is the RTX 4090 with the budget picks being either a used RTX 3090 or a used RTX 3090 Ti.
- kristianp 3y agoFor a compromise, how is the recently released 4060ti with 16gb RAM? Its about a third the price of a 4090.
- Tepix 3y agoI used Tim's guide to build a dual RTX 3090 PC, paying 2300€ in total by getting used components. It can run inference of Llama-65B 4bit quantized at more than 10tok/s. Specs: 2x RTX 3090, NVLink Bridge, 128GB DDR4 3200 RAM, Ryzen 7 3700X, X570 SLI mainboard, 2TB M.2 NVMe SSD, air cooled mesh case. Finding the 3-slot nvlink bridge is hard and it's usually expensive. I think it's not worth it in most cases. I managed to find a cheap used one. Cooling is also a challenge. The cards are 2.7 slots wide and the spacing is usually 3 slots, so there isn't much room. Some people are putting 3d printed shrouds on the back of the PC case to suck the air out of the cards with an extra external fan. Also limiting the power from 350W to 280W or so per card doesn't cost a lot of performance. The CPU is not limiting the performance at all, as long as you have 4 cores per GPU you're good.
- bwv848 3y agoManaged to snatch a 3090 during the GPU shortage in 2020. Did a lot of training and mining, and got some of my results published, think I gained much more than the cost of the hardware purchases. Kinda miss the day of eth mining. 3090 is a still good card and I'm pretty sure your rig is going to serve you well. ps: ~280W power limit is a good call, it won't heat up your room too much.
- horsawlarway 3y agoMy build is close to this. I purchased everything new except the 3090s, and I paid about $3000. 2x RTX 3090 128 GB DDR5 Intel core i9 600 series Z790 Mainboard I used Intel instead of AMD for the cpu, which pushed my prices higher... but I saved on the back side by skipping the NVLink Bridge. Good to know I'm not missing much with out the Bridge, since I get about 13tok/s on Llama-65B 4 bit if I push all layers onto the GPU.
- frognumber 3y agoI think there's one more axis: Frequency-of-use. For occasionally use, the major constraint isn't speed so much as which models fit. I tend to look at $/GB VRAM as my major spec. Something like a 3060 12GB is an outlier for fitting sensible models while being cheap. I don't mind waiting a minute instead of 15 seconds for some complex inference if I do it a few times per day. Or having training be slower if it comes up once every few months.
- _cnmh 3y agoHopefully the next generation of cards have high-VRAM variants.
- bick_nyers 3y ago"As for capacity, Samsung’s first GDDR7 chips are 16Gb, matching the existing density of today’s top GDDR6(X) chips. So memory capacities on final products will not be significantly different from today’s products, assuming identical memory bus widths. DRAM density growth as a whole has been slowing over the years due to scaling issues, and GDDR7 will not be immune to that." Source: https://www.anandtech.com/show/18963/samsung-completes-initial-gddr7-development-first-parts-to-reach-up-to-32gbpspin https://www.anandtech.com/show/18963/samsung-completes-initi...
- frognumber 3y agoI can buy a DDR5 64GB kit from Crucial for $160. https://www.crucial.com/memory/ddr5/ct2k32g48c40u5 https://www.crucial.com/memory/ddr5/ct2k32g48c40u5 If a $1000 GPU came with that, it would blow everything else out-of-the-water for model size. Speed? No. Model size? Yes. If it came with 320GB, I could run ChatGPT-grade LLMs. That's $800 worth of DDR5. Instead, I get 24GB on the 3090 or 4090 for $2k. A $3k LLM-capable card would not be a hard expense to justify.
- lyapunova 3y agoI never tire of this. Tim is a wonderful no nonsense person. I love these posts and I love that it stays up to date.
- jcuenod 3y agoAny advice for mobile gpus? I'm interested in getting a laptop (preferably in the portable category). Obviously it's not going to be in 4090 territory, that's a tradeoff I'm willing to make.
- paul_funyun 3y agoOne, don't use a case. Look at how miners mounted their hardware on racks and take notes. Cheaper, better for temps, and the most efficient use of space. Two, I recommend ignoring electricity cost and using all you can. If it's cheaper now than it ever will be, use it while it's cheap. If it will go down due to renewables, nuclear, etc in the future, it's good to buy up the GPUs while their price is artificially depressed from energy fears. Third, go for server type PSUs and breakout boards. The server PSUs cant be beaten in watts for your dollar, and are extremely efficient. Finally, consider scooping up some x79 and x99 xeon boards from Chinese sellers. They're cheap as hell, have PCI lanes out the wazoo, etc. This means you don't have to fool with as many mobos to run the same amount of gpus. If you go this route, don't get the bottom of the barrel no-name motherboards. Machinist is a decent one.
- justinclift 3y agoRaw performance rating for the RTX 3070 seems very weirdly placed in the chart. It's below the RTX 3060 Ti, which doesn't seem to make any sense.
- adultSwim 3y agoWeird to leave out Apple. They seem to be the cheapest option to get a large amount of GPU memory.