3 ms·
Considering how Nvidia uses GPU RAM to segment their products between high end consumer GPUs (where the 4090 hasn't progressed at all from the 3090 in still hav
by ahefner 4y ago
Considering how Nvidia uses GPU RAM to segment their products between high end consumer GPUs (where the 4090 hasn't progressed at all from the 3090 in still having 24 GB RAM) versus the crazy high end ML accelerators, I've recently wondered how it would work out for AMD if they slapped 48+ GB RAM (sufficient to run GPT-NeoX-20B..) on one of their higher end parts (ideally without a large price premium) in order to excite the people trying to run ML on the cheap that are currently bottlenecked by RAM on consumer cards and spur the software development required to get more things running on AMD.
- kkielhofner 4y agoLike an AMD version of an RTX A6000? Granted $4500 still isn’t “cheap” but for a dual slot blower card with great performance and 48GB they do really well in the mid-range of ML. I don’t really pay attention to AMD cards anymore. I’m not a gamer and every time over the past five years I’ve put a toe in the water of the disaster that is ROCm, their drivers, “support” in frameworks, etc I run back to CUDA. It really seems like AMD has ceded everything other than supercomputer cluster ML to Nvidia and CUDA. They do well in gaming and they can sell and support the extremely high end of ML with mostly bespoke applications that have teams of professional support and institutional backing. When you’re spending $100 million or whatever on hardware flop/$ really counts and you can throw people at making it work. In the low to mid ranges of ML the 25% or even 50% savings on hardware is quickly lost when compared to the human time sink that is the entire AMD ML “ecosystem”. If you’re buying 10 GPUs you’re talking about saving what, $20K? That “savings” gets wrecked really quickly when you look at salaries and time spent chasing yet another bug/quirk/showstopper with ROCm.
- ahefner 4y agoYeah, I suppose the "run it local and on a budget" crowd isn't that big, but maybe they do have some free time to improve the software ecosystem.
- kkielhofner 4y agoThe problem is “on a budget” applies if your time is free. Most of the ML ecosystem is still gold plated - talent, hardware, etc with talent being the biggest cost for most applications. Frankly, Nvidia doesn’t exactly foster goodwill (just look at any HN thread when they come up). I suspect that a decent number of people are willing to throw time and pull requests into the ROCm ecosystem just out of spite for Nvidia. Good for them because Nvidia needs a serious competitor! I’m rooting for AMD but I’ve always been in the position of “it needs to work now in the shortest and cheapest path of least resistance” which has always been, and will be for the foreseeable future, CUDA.
- nerdyadventurer 4y agoI am also really passionate to see AMD cards for consumer ML workloads.