5 ms·
AMD GPUs are becoming a serious contender for LLM inference. vLLM is already showing impressive performance on AMD [1], even with consumer-grade Radeon cards (e
by pinsiang 2y ago
AMD GPUs are becoming a serious contender for LLM inference. vLLM is already showing impressive performance on AMD [1], even with consumer-grade Radeon cards (even support GGUF) [2]. This could be a game-changer for folks who want to run LLMs without shelling out for expensive NVIDIA hardware.
[1] https://blog.vllm.ai/2024/10/23/vllm-serving-amd.html https://blog.vllm.ai/2024/10/23/vllm-serving-amd.html
[2] https://embeddedllm.com/blog/vllm-now-supports-running-gguf-on-amd-radeon-gpu https://embeddedllm.com/blog/vllm-now-supports-running-gguf-...
- MrBuddyCasino 2y agoFun fact: Nvidia H200 are currently half the price/hr of H100 bc people can’t get vLLM to work on it. https://x.com/nisten/status/1871325538335486049 https://x.com/nisten/status/1871325538335486049
- treprinum 2y agoAMD decided not to release a high-end GPU this cycle so any investment into 7x00 or 6x00 is going to be wasted as Nvidia 5x00 is likely going to destroy any ROI from the older cards and AMD won't have an answer for at least two years, possibly never due to being non-existing in high-end consumer GPUs usable for compute.
- BearOso 2y agoNo high-end consumer RDNA4 GPU this cycle. And it's only missing the very high-end model. So we'll still get at least a 7800xt equivalent and whatever CDNA MI models they come out with. The market for the extreme high-end consumer is pretty small, so they're only missing out on clout.
- treprinum 2y agoThe top-end RDNA4 GPU will have 16GB RAM. That's a massive regression compared to 7900XTX and performance-wise it should be at best at the 7900XTX level. We are discussing AMD cards for LLM inference where VRAM is arguably the most important aspect of a GPU and AMD just threw in the towel for this cycle.
- latchkey 2y agoThese blog posts were written based on my company, Hot Aisle, donating the compute. =) Super proud of being able to support this.