5 ms·
What GPU offers a good balance between cost and performance for running LLMs locally? I'd like to do more experimenting, and am due for a GPU upgrade from my 10
by dividefuel 2y ago
What GPU offers a good balance between cost and performance for running LLMs locally? I'd like to do more experimenting, and am due for a GPU upgrade from my 1080 anyway, but would like to spend less than $1600...
- redacted 2y agoNvidia for compatibility, and as much VRAM as you can afford. Shouldn't be hard to find a 3090 / Ti in your price range. I have had decent success with a base 3080 but the 10GB really limits the models you can run
- yk 2y agoSecond the other comment, as much vram as possible. A 3060 has 12 GB at a reasonable price point. (And is not too limiting.)
- grobbyy 2y agoThere's a huge step to I'm capability with 16gb and 24gb, for not to much more. The 4060 has a 16gb version, for example. On the the cheap end, the Intel Arc does too. Next major step up is 48GB and then hundreds of GB. But a lot of ML models target 16-24gb since that's in the grad student price range.
- navbaker 2y agoAt the 48GB level, L40S are great cards and very cost effective. If you aren’t aiming for constant uptime on several >70B models at once, they’re for sure the way to go!
- nickthegreek 2y agoDual 3090s are way cheaper than the l40s though. You can even buy a few backups.
- navbaker 2y agoYeah, I’m specifically responding to the parent’s comment about the 48GB tier. When you’re looking in that range, it’s usually because you want to pack in as much vram as possible into your rack space, so consumer level cards are off the table. I definitely agree multiple 3090 is the way to go if you aren’t trying to host models for smaller scale enterprise use, which is where 48GB cards shine.
- bubaumba 2y ago> L40S are great cards and very cost effective from https://www.asacomputers.com/nvidia-l40s-48gb-graphics-card.html https://www.asacomputers.com/nvidia-l40s-48gb-graphics-card.... nvidia l40s 48gb graphics card Our price: $7,569.10* Not arguing against 'great', but cost efficiency is questionable. for 10% you can get two used 3090. The good thing about LLMs is they are sequential and should be easily parallelized. Model can be split in several sub-models, by the number of GPUs. Then 2,3,4.. GPUs should improve performance proportionally on big batches, and make it possible to run bigger model on low end hardware.
- bloomingkales 2y agoHonestly I think if you just want to do inferencing the 7600xt and rx6800 have 16gb at $300 and $400 on Amazon. It's gonna be my stop gap until whatever. The RX6800 has better memory bandwidth than the 4060ti (think it matches the 4070).
- kolbe 2y agoAMD GPUs are a fantastic deal until you hit a problem. Some models/frameworks it works great. Others, not so much.
- bloomingkales 2y agoFor sure but I think people on the fine tuning/training/stable diffusion side are more concerned with that. They make a big fuss about this and basically talk people out of a perfectly good and well priced 16gb vram card that literally works out of the box with ollama, lmstudio for text inferencing. Kind of one of the reasons AMD is a sleeper stock for me. If people only knew.
- kolbe 2y agoIf you want to wait until the 5090s come out, you should see a drop in the price of the 30xx and 40xx series. Right now, shopping used, you can get two 3090s or two 4080s in your price range. Conventional wisdom says two 3090s would be better, but this is all highly dependent on what models you want to run. Basically the first requirement is to have enough VRAM to host all of your model on it, and secondarily, the quality of the GPU. Have a look through Hugging Face to see which models interest you. A rough estimate for the amount of VRAM you need is half the model size plus a couple gigs. So, if using the 70B models interests you, two 4080s wouldn't fit it, but two 3090s would. If you're just interested in the 1B, 3B and 7B models (llama 3B is fantastic), you really don't need much at all. A single 3060 can handle that, and those are not expensive.
- zitterbewegung 2y agoGet a new / used 3090 it has 24GB of RAM and it's below $1600.
- christianqchung 2y agoA lot of moderate power users are running an undervolted used pair of 3090s on a 1000-1200W psu. 48 GB of vram let's you run 70B models at Q4 with 16k context. If you use speculative decoding (a small model generates tokens verified by a larger model, I'm not sure on the specifics) you can get past 20 tokens per second it seems. You can also fit 32B models like Qwen/Qwen Coder at Q6 with lots of context this way, with spec decoding, closer to 40+ tks/s.
- adam_arthur 2y agoInferencing does not require Nvidia GPUs at all, and its almost criminal to be recommending dedicated GPUs with only 12GB of RAM. Buy a MacMini or MacbookPro with RAM maxed out. I just bought an M4 mac mini for exactly this use case that has 64GB for ~2k. You can get 128GB on the MBP for ~5k. These will run much larger (and more useful) models. EDIT: Since the request was for < $1600, you can still get a 32GB mac mini for $1200 or 24GB for $800
- natch 2y agoReasonable? $7,000 for a laptop is pretty up there. [Edit: OK I see I am adding cost when checking due to choosing a larger SSD drive, so $5,000 is more of a fair bottom price, with 1TB of storage.] Responding specifically to this very specific claim: "Can get 128GB of ram for a reasonable price." I'm open to your explanation of how this is reasonable — I mean, you didn't say cheap, to be fair. Maybe 128GB of ram on GPUs would be way more (that's like 6 x 4090s), is what you're saying. For anyone who wants to reply with other amounts of memory, that's not what I'm talking about here. But on another point, do you think the ram really buys you the equivalent of GPU memory? Is Apple's melding of CPU/GPU really that good? I'm not just coming from a point of skepticism, I'm actually kind of hoping to be convinced you're right, so wanting to hear the argument in more detail.
- adam_arthur 2y agoIt's reasonable in a "working professional who gets substantial value from" or "building an LLM driven startup project" kind of way. It's not for the casual user, but for somebody who derives significant value from running it locally. Personally I use the MacMini as a hub for a project I'm working on as it gives me full control and is simply much cheaper operationally. A one time ~$2000 cost isn't so bad for replacing tasks that a human would have to do. e.g. In my case I'm parsing loosely organized financial documents where structured data isn't available. I suspect the hardware costs will continue to decline rapidly as they have in the past though, so that $5k for 128GB will likely be $5k for 256GB in a year or two, and so on. We're almost at the inflection point where really powerful models are able to be inferenced locally for cheap
- throawayonthe 2y ago[dead]
- elorant 2y agoI consider the RTX 4060 Ti as the best entry level GPU for running small models. It has 16GBs of RAM which gives you plenty of space for running large context windows and Tensor Cores which are crucial for inference. For larger models probably multiple RTX 3090s since you can buy them on the cheap on the second hand market. I don’t have experience with AMD cards so I can’t vouch for them.
- fnqi8ckfek 2y agoI know nothing about gpus. Should I be assuming that when people say "ram" in the context of gpus they always mean vram?
- thijson 2y agoI've been running some of the larger models (like Llama 405B) via CPU on a Dell R820. It's got 32 Xeon's (4 chips), and 256 GB RAM. I bought it used for around $400. The memory is NUMA, so it makes sense if the computing is done on local data, not sure if Ollama supports that. The tokens per second is very slow though, but at least it can execute it. I think the future will be increasingly more powerful NPU's built into future CPU's. That will need to paired with higher bandwidth memory, maybe HBM, or silicon photonics for off chip memory.