3 ms·
Our team has access to multiple systems that either have MI250Xs or H100s. Getting stuff to work with AMD/ROCm is substantially more effort than the NVIDIA/CUDA
by thomasfedb 2y ago
Our team has access to multiple systems that either have MI250Xs or H100s. Getting stuff to work with AMD/ROCm is substantially more effort than the NVIDIA/CUDA experience.
Some of this is lack of groundwork/engineering by packages or system administrators, but it seems a decent amount is the relative lack of effort by AMD to make things work well OOTB.
- BoingBoomTschak 2y agoThe real question is: is this from lack of effort or simply from NVidia's headstart? Will it get better?
- noch 2y ago> Will it get better? It won't, not in any way that will make AMD approximately competitive with Nvidia. AMD, unlike Nvidia, seems unable to prioritize developers. Here's a summary of last week's charlie-fox when the TinyGrad team attempted to get 2 MI300s for on-premises testing and was rebuffed by an AMD representative. https://x.com/dehypokriet/status/1879974587082912235 https://x.com/dehypokriet/status/1879974587082912235
- musicale 2y agoBoth, really. It will get better, but Nvidia seems to be creating CUDA libraries for all kinds of applications, so the moat is constantly widening/deepening.
- latchkey 2y agoInstalling ROCm is easy and well documented [0]. Anush (AMD VP of AI software) has had a fire lit under his butt after the recent SemiAnalysis article [1] and is actively taking feedback on improving the experience. If you have specific things you'd like to see, I'm more than happy to forward them onto him (contact in my profile). [0] https://rocm.docs.amd.com/en/latest/ https://rocm.docs.amd.com/en/latest/ [1] https://semianalysis.com/2024/12/22/mi300x-vs-h100-vs-h200-benchmark-part-1-training/ https://semianalysis.com/2024/12/22/mi300x-vs-h100-vs-h200-b...