12 ms·
Making AMD GPUs competitive for LLM inference (2023)
- dragontamer 2y agoIntriguing. I thought AMD GPUs didn't have tensor cores (or matrix multiplication units) like NVidia. I believe they are only dot product / fused multiply and accumulate instructions. Are these LLMs just absurdly memory bound so it doesn't matter?
- ryao 2y agoThey don’t, but GPUs were designed for doing matrix multiplications even without the special hardware instructions for doing matrix multiplication tiles. Also, the forward pass for transformers is memory bound, and that is what does token generation.
- dragontamer 2y agoWell sure, but in other GPU tasks, like Raytracing, the difference between these GPUs is far more pronounced. And AMD has passable Raytracing units (NVidias are better but the difference is bigger than these LLM results). If RAM is the main bottleneck then CPUs should be on the table.
- webmaven 2y agoRAM is (often) the bottleneck for highly parallel GPUs, but not for CPUs. Though the distinction between the two categories is blurring.
- ryao 2y agoMemory bandwidth is the bottleneck for both when running GEMV, which is the main operation used by token generation in inference. It has always been this way.
- IX-103 2y ago> If RAM is the main bottleneck then CPUs should be on the table That's certainly not the case. The graphics memory model is very different from the CPU memory model. Graphics memory is explicitly designed for multiple simultaneous reads (spread across several different buses) at the cost of generality (only portions of memory may be available on each bus) and speed (the extra complexity means reads are slower). This makes then fast at doing simple operations on a large amount of data. CPU memory only has one bus, so only a single read can happen at a time (a cache line read), but can happen relatively quickly. So CPUs are better for workloads with high memory locality and frequent reuse of memory locations (as is common in procedural programs).
- dragontamer 2y ago> CPU memory only has one bus If people are paying $15,000 or more per GPU, then I can choose $15,000 CPUs like EPYC that have 12-channels or dual-socket 24-channel RAM. Even desktop CPUs are dual-channel at a minimum, and arguably DDR5 is closer to 2 or 4 buses per channel. Now yes, GPU RAM can be faster, but guess what? https://www.tomshardware.com/pc-components/cpus/amd-crafts-custom-epyc-cpu-for-microsoft-azure-with-hbm3-memory-cpu-with-88-zen-4-cores-and-450gb-of-hbm3-may-be-repurposed-mi300c-four-chips-hit-7-tb-s https://www.tomshardware.com/pc-components/cpus/amd-crafts-c... GPUs are about extremely parallel performance, above and beyond what traditional single-threaded (or limited-SIMD) CPUs can do. But if you're waiting on RAM anyway?? Then the compute-method doesn't matter. Its all about RAM.
- ryao 2y agoWhere are these GPUs with multiple buses? I only know of GPUs with wide buses.
- schmidtleonard 2y agoCPUs have pitiful RAM bandwidth compared to GPUs. The speeds aren't so different but GPU RAM busses are wiiiiiiiide.
- teleforce 2y agoCompute Express Link (CXL) should mostly solve limited RAM with CPU: 1) Compute Express Link (CXL): https://en.wikipedia.org/wiki/Compute_Express_Link https://en.wikipedia.org/wiki/Compute_Express_Link PCIe vs. CXL for Memory and Storage: https://news.ycombinator.com/item?id=38125885 https://news.ycombinator.com/item?id=38125885
- schmidtleonard 2y agoGigabytes per second? What is this, bandwidth for ants? My years old pleb tier non-HBM GPU has more than 4 times the bandwidth you would get from a PCIe Gen 7 x16 link, which doesn't even officially exist yet.
- teleforce 2y agoYes CXL will soon benefit from PCIe Gen 7 x16 with expected 64GB/s in 2025 and the non-HBM bandwidth I/O alternative is increasing rapidly by the day. For most inferences of near real-time LLM it will be feasible. For majority of SME companies and other DIY users (humans or ants) with their localized LLM should not be any issues [1],[2]. In addition new techniques for more efficient LLM are being discover to reduce the memory consumption [3]. [1] Forget ChatGPT: why researchers now run small AIs on their laptops: https://news.ycombinator.com/item?id=41609393 https://news.ycombinator.com/item?id=41609393 [2] Welcome to LLMflation – LLM inference cost is going down fast: https://a16z.com/llmflation-llm-inference-cost/ https://a16z.com/llmflation-llm-inference-cost/ [3] New LLM optimization technique slashes memory costs up to 75%: https://news.ycombinator.com/item?id=42411409 https://news.ycombinator.com/item?id=42411409
- schmidtleonard 2y ago
- ryao 2y agoThe main bottleneck is memory bandwidth. CPUs have less memory bandwidth than GPUs.
- throwaway314155 2y ago> Are these LLMs just absurdly memory bound so it doesn't matter? During inference? Definitely. Training is another story.
- boroboro4 2y agoThey absolutely do have similar cores to tensor cores, it's called matrix cores. And they have particular instructions to utilize them (MFMA). Note I'm talking about DC compute chips, like MI300. LLMs aren't memory bound in production loads, they are pretty much compute bound too, at least in prefill phase, but in practice in general too.
- almostgotcaught 2y agoYa people in these comments don't know what they're talking about (no one ever does in these threads). AMDGPU has had MMA and WMMA for a while now https://rocm.docs.amd.com/projects/rocWMMA/en/latest/what-is-rocwmma.html https://rocm.docs.amd.com/projects/rocWMMA/en/latest/what-is...
- sroussey 2y ago[2023] Btw, this is from MLC-LLM which makes WebLLM and other good stuff.
- throwaway314155 2y ago> Aug 9, 2023 Ignoring the very old (in ML time) date of the article... What's the catch? People are still struggling with this a year later so I have to assume it doesn't work as well as claimed. I'm guessing this is buggy in practice and only works for the HF models they chose to test with?
- deleted 2y ago[deleted]
- Const-me 2y agoIt’s not terribly hard to port ML inference to alternative GPU APIs. I did it for D3D11 and the performance is pretty good too: https://github.com/Const-me/Cgml https://github.com/Const-me/Cgml The only catch is, for some reason developers of ML libraries like PyTorch aren’t interested in open GPU APIs like D3D or Vulkan. Instead, they focus on proprietary ones i.e. CUDA and to lesser extent ROCm. I don’t know why that is. D3D-based videogames are heavily using GPU compute for more than a decade now. Since Valve shipped SteamDeck, the same now applies to Vulkan on Linux. By now, both technologies are stable, reliable and performant.
- jsheard 2y agoIsn't part of it because the first-party libraries like cuDNN are only available through CUDA? Nvidia has poured a ton of effort into tuning those libraries so it's hard to justify not using them.
- Const-me 2y agoUnlike training, ML inference is almost always bound by memory bandwidth as opposed to computations. For this reason, tensor cores, cuDNN, and other advanced shenanigans make very little sense for the use case. OTOH, general-purpose compute instead of fixed-function blocks used by cuDNN enables custom compression algorithms for these weights which does help, by saving memory bandwidth. For example, I did custom 5 bits/weight quantization which works on all GPUs, no hardware support necessary, just simple HLSL codes: https://github.com/Const-me/Cgml?tab=readme-ov-file#bcml1-codec https://github.com/Const-me/Cgml?tab=readme-ov-file#bcml1-co...
- shihab 2y agoI have come across quite few startups who are trying a similar idea: break the nvidia monopoly by utilizing AMD GPUs (for inference at least): Felafax, Lamini, tensorwave (partially), SlashML. Even saw optimistic claims like CUDA moat is only 18 months deep from some of them [1]. Let's see. [1] https://www.linkedin.com/feed/update/urn:li:activity:7275885292513906689/ https://www.linkedin.com/feed/update/urn:li:activity:7275885...
- ryukoposting 2y agoPeculiar business model, at a glance. It seems like they're doing work that AMD ought to be doing, and is probably doing behind the scenes. Who is the customer for a third-party GPU driver shim?
- tesch1 2y agoAMD. Just one more dot to connect ;)
- dpkirchner 2y agoCould be trying to make themselves a target for a big acquihire.
- dboreham 2y ago> Could be trying to make themselves a target for a big acquihire. Is this something anyone sets out to do?
- seeknotfind 2y agoYes.
- ryukoposting 2y agoIt definitely is, yes.
- to11mtm 2y ago
- jroesch 2y agoNote: this is old work, and much of the team working on TVM, and MLC were from OctoAI and we have all recently joined NVIDIA.
- sebmellen 2y agoIs there no hope for AMD anymore? After George Hotz/Tinygrad gave up on AMD I feel there’s no realistic chance of using their chips to break the CUDA dominance.
- llm_trw 2y agoNot really. AMD is constitutionally incapable of shipping anything but mid range hardware that requires no innovation. The only reason why they are doing so well in CPUs right now is that Intel has basically destroyed itself without any outside help.
- perching_aix 2y agoAnd I'm supposed to believe that HN is this amazing platform for technology and science discussions, totally unlike its peers...
- zamadatix 2y agoThe above take is worded a bit cynical but is their general approach to GPUs lately across the board e.g. https://www.techpowerup.com/326415/amd-confirms-retreat-from-the-enthusiast-gpu-segment-to-focus-on-gaining-market-share https://www.techpowerup.com/326415/amd-confirms-retreat-from... Also I'd take HN as being being an amazing platform for the overall consistency and quality of moderation. Anything beyond that depends more on who you're talking to than where at.
- petesergeant 2y agoMaybe be the change you want to see and tell us what the real story is?
- lasermike026 2y agoI believe these efforts are very important. If we want this stuff to be practical we are going to have to work on efficiency. Price efficiency is good. Power and compute efficiency would be better. I have been playing with llama.cpp to run interference on conventional cpus. No conclusions but it's interesting. I need to look at llamafile next.
- zamalek 2y agoI have been playing around with Phi-4 Q6 on my 7950x and 7900XT (with HSA_OVERRIDE_GFX_VERSION). It's bloody fast, even with CPU alone - in practical terms it beats hosted models due to the roundtrip time. Obviously perf is more important if you're hosting this stuff, but we've definitely reached AMD usability at home.
- slavik81 2y agoIf you're not using your iGPU, you can disable it in BIOS and you won't need to set HSA_OVERRIDE_GFX_VERSION.
- latchkey 2y agoPreviously: Making AMD GPUs competitive for LLM inference https://news.ycombinator.com/item?id=37066522 https://news.ycombinator.com/item?id=37066522 (August 9, 2023 — 354 points, 132 comments)
- lxe 2y agoA used 3090 is $600-900, performs better than 7900, and is much more versatile because CUDA
- Uehreka 2y agoReality check for anyone considering this: I just got a used 3090 for $900 last month. It works great. I would not recommend buying one for $600, it probably either won’t arrive or will be broken. Someone will reply saying they got one for $600 and it works, that doesn’t mean it will happen if you do it. I’d say the market is realistically $900-1100, maybe $800 if you know the person or can watch the card running first. All that said, this advice will expire in a month or two when the 5090 comes out.
- idonotknowwhy 2y agoI've bought 5 used and they're all perfect. But that's what buyer protection on ebay is for. Had to send back an Epyc mobo with bent pins and ebay handled it fine.
- fireant 2y agoI've bought used 3090 last year for ML and while it works fine, has correct DRAM and stuff, when I tried gaming on it I've noticed that it is significantly slower than my 3080. I'm not sure if the seller has pulled some shenanigans on me or the card actually degraded during whatever mining they did. Just beware, the card might be "working fine" on a first glance, but actually be damaged.
- ryao 2y agoI got a refurbished $800 3090 Ti FE earlier this year from microcenter. Sadly, they sold out and never restocked.
- coolspot 2y agoZotac official website has refurb 3090 ti for $899
- nullbyte808 2y agoThis benchmark doest look right. Is it using the tensor cores in the Nvidia gpu? AMD does not have AI cores so should run noticeably slower.
- nomel 2y agoAMD has WMMA.
- deleted 2y ago[deleted]
- mattfrommars 2y agoGreat, I have yet to understand why does not the ML community really push or move away from CUDA? To me, it feel like a dinosaur move to build on top of CUDA which is screaming proprietary nothing about it is open source or cross platform. The reason why I say its dinosaur is, imagine, we as a dev community continued to build on top of Flash or Microsoft Silverlight... LLM and ML has been out for quiet a while, with AI/LLM advancement, the transition must have been much quicker to move cross platform. But this hasn't yet and not sure when it will happen. Building a translation layer on top CUDA is not the answer either to this problem.
- dwood_dev 2y agoExcept I never hear complaints about CUDA from a quality perspective. The complaints are always about lock in to the best GPUs on the market. The desire to shift away is to make cheaper hardware with inferior software quality more usable. Flash was an abomination, CUDA is not.
- xedrac 2y agoMaybe the situation has gotten better in recent years, but my experience with Nvidia toolchains was a complete nightmare back in 2018.
- claytonjy 2y agoThe cuda situation is definitely better. The nvidia struggles are now with the higher-level software they’re pushing (triton, tensor-llm, riva, etc), tools that are the most performant option when they work, but a garbage developer experience when you step outside the golden path
- cameron_b 2y agoI want to double-down on this statement, and call attention to the competitive nature of it. Specifically, I have recently tried to set up Triton on arm hardware. One might presume Nvidia would give attention to an architecture they develop, but the way forward is not easy. For some version of Ubuntu, you might have the correct version of python ( usually older than packaged ) but current LTS is out of luck for guidance or packages. https://github.com/triton-lang/triton/issues/4978 https://github.com/triton-lang/triton/issues/4978
- pavelstoev 2y agoThe problem is that performance achievements on AMD consumer-grade GPUs (RX7900XTX) are not representative/transferrable to the Datacenter grade GPUs (MI300X). Consumer GPUs are based on RDNA architecture, while datacenter GPUs are based on the CDNA architecture, and only sometime in ~2026 AMD is expected to release unifying UDNA architecture [1]. At CentML we are currently working on integrating AMD CDNA and HIP support into our Hidet deep learning compiler [2], which will also power inference workloads for all Nvidia GPUs, AMD GPUs, Google TPU and AWS Inf2 chips on our platform [3] [1] https://www.jonpeddie.com/news/amd-to-integrate-cdna-and-rdna-architectures-to-compete-in-ai/#:~:text=In%202019%2C%20AMD%20transitioned%20away,Well%2C%20there's%20plenty%20of%20time https://www.jonpeddie.com/news/amd-to-integrate-cdna-and-rdn.... [2] https://centml.ai/hidet/ https://centml.ai/hidet/ [3] https://centml.ai/platform/ https://centml.ai/platform/
- llm_trw 2y agoThe problem is that the specs of AMD consumer-grade GPUs do not translate to computer performance when you try and chain more than one together. I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.
- tspng 2y agoWow, are these 7 RTX 4090s in a single setup? Care to share more how you build it (case, cooling, power, ..)?
- aussieguy1234 2y agoI got a "gaming" PC for LLM inference with an RTX 3060. I could have gotten more VRAM for my buck with AMD, but didn't because at the time alot of inference needed CUDA. As soon AMD is as good as Nvidia for inference, I'll switch over. But I've read on here that their hardware engineers aren't even given enough hardware to test with...
- lhl 2y agoJust an FYI, this is writeup from August 2023 and a lot has changed (for the better!) for RDNA3 AI/ML support. That being said, I did some very recent inference testing on an W7900 (using the same testing methodology used by Embedded LLM's recent post to compare to vLLM's recently added Radeon GGUF support [1]) and MLC continues to perform quite well. On Llama 3.1 8B, MLC's q4f16_1 (4.21MB weights) performed +35% faster than llama.cpp w/ Q4_K_M w/ their ROCm/HIP backend (4.30MB weights, 2% size difference). That makes MLC still the generally fastest standalone inference engine for RDNA3 by a country mile. However, you have much less flexibility with quants and by and large have to compile your own for every model, so llama.cpp is probably still more flexible for general use. Also llama.cpp's (recently added to llama-server) speculative decoding can also give some pretty sizable performance gains. Using a 70B Q4_K_M + 1B Q8_0 draft model improves output token throughput by 59% on the same ShareGPT testing. I've also been running tests with Qwen2.5-Coder and using a 0.5-3B draft model for speculative decoding gives even bigger gains on average (depends highly on acceptance rate). Note, I think for local use, vLLM GGUF is still not suitable at all. When testing w/ a 70B Q4_K_M model (only 40GB), loading, engine warmup, and graph compilation took on avg 40 minutes. llama.cpp takes 7-8s to load the same model. At this point for RDNA3, basically everything I need works/runs for my use cases (primarily LLM development and local inferencing), but almost always slower than an RTX 3090/A6000 Ampere (a new 24GB 7900 XTX is $850 atm, used or refurbished 24 GB RTX 3090s are in in the same ballpark, about $800 atm; a new 48GB W7900 goes for $3600 while an 48GB A6000 (Ampere) goes for $4600). The efficiency gains can be sizable. Eg, on my standard llama-bench test w/ llama2-7b-q4_0, the RTX 3090 gets a tg128 of 168 t/s while the 7900 XTX only gets 118 t/s even though both have similar memory bandwidth (936.2 GB/s vs 960 GB/s). It's also worth noting that since the beginning of the year, the llama.cpp CUDA implementation has gotten almost 25% faster, while the ROCm version's performance has stayed static. There is an actively (solo dev) maintained fork of llama.cpp that sticks close to HEAD but basically applies a rocWMMA patch that can improve performance if you use the llama.cpp FA (still performs worse than w/ FA disabled) and in certain long-context inference generations (on llama-bench and w/ this ShareGPT serving test you won't see much difference) here: https://github.com/hjc4869/llama.cpp https://github.com/hjc4869/llama.cpp - The fact that no one from AMD has shown any interest in helping improve llama.cpp performance (despite often citing llama.cpp-based apps in marketing/blog posts, etc is disappointing ... but sadly on brand for AMD GPUs). Anyway, for those interested in more information and testing for AI/ML setup for RDNA3 (and AMD ROCm in general), I keep a doc with lots of details here: https://llm-tracker.info/howto/AMD-GPUs https://llm-tracker.info/howto/AMD-GPUs [1] https://embeddedllm.com/blog/vllm-now-supports-running-gguf-on-amd-radeon-gpu https://embeddedllm.com/blog/vllm-now-supports-running-gguf-...
- mrcsharp 2y agoI will only consider AMD GPUs for LLM when I can easily make my AMD GPU available within WSL and Docker on Windows. For now, it is as if AMD does not exist in this field for me.
- e-max 2y agoIsn't it already available somehow? I didn't test it seriously, I just needed to quickly run Whisper but $ rocminfo | grep -E "WSL|XTX" WSL environment detected. Marketing Name: AMD Radeon RX 7900 XTX
- mrcsharp 2y agoInteresting. Looking here [1] it seems like this is a thing now. Will dive deeper into this later but looks promising. [1] https://rocm.docs.amd.com/projects/radeon/en/latest/docs/install/wsl/install-radeon.html https://rocm.docs.amd.com/projects/radeon/en/latest/docs/ins...
- Sparkyte 2y agoMore players in the market the better. AI shouldn't be owned by one business.
- starlite-5008 2y ago[dead]
- melodyogonna 2y agoModular claims that it achieves 93% GPU utilization on AMD GPUs [1], official preview release coming early next year, we'll see. I must say I'm bullish because of feedback I've seen people give about the performance on Nvidia GPUs 1.https://www.modular.com/max https://www.modular.com/max
- guerrilla 2y agoSo, does ollama use this work or does it do something else? How does it compare?
- varelse 2y ago[dead]