3 ms·
Christ GPU prices have gotten crazy How do AMD cards perform with LLMs? A 9070 is sold for ~$600 and has 16GB VRAM
by tripleee 4mo ago
Christ GPU prices have gotten crazy
How do AMD cards perform with LLMs? A 9070 is sold for ~$600 and has 16GB VRAM
- lambda 4mo agoThat should do pretty well. Memory bandwidth is the biggest bottleneck for token generation, at 644 GB/s you should be able to do pretty well on a 9070, while prompt proessing is more compute bound and Nvidia tends to have the edge there. 16 GiB won't fit you much, so you'd probably want at least 2x, and preferably 3x of those, and then you need a motherboard, power, etc. that can handle that.
- overgard 4mo agoIn my personal experience, I wouldn't bother with 16GB cards for coding -- the useful models are _slightly_ too large to work at any reasonable speed
- toyg 4mo agoThat's not my experience, and the trajectory is good anyway - what doesn't work perfectly today will be just fine in a few months. In a quickly moving field, it's amazing how much money one can save by overcoming FOMO and not living on the bleeding edge. It's like waiting for Steam sales, the games will be just as good.
- overgard 4mo agoCurious what model you're using that works well on a 16GB card? I very much want to use my 5080 for inference, but everything I've tried so far has either just not been good enough or painfully slow.
- toyg 4mo agoQwen 3.5-9b-Q4_K_M. I have a 5080 too! For me, the key has been dropping Ollama for Llama.cpp, which is not particularly scary to configure anymore and just skyrocketed performance. I download the models with LM Studio, then run them with llama.cpp.
- brianvoe 4mo agogemma 4 12b
- gargola_ 4mo agoThis can work for some things but I wanted to run hermes with 16gb but 12b is too low and the context was too limited, they recommend 27b