4 ms·
As more a business person than engineer, help me understand why AMD are not getting this, what's the counter argument? Is CUDA just too far ahead, are they lack
by monkeydust 2y ago
As more a business person than engineer, help me understand why AMD are not getting this, what's the counter argument? Is CUDA just too far ahead, are they lacking the right people in senior leadership roles to see this through?
- cyanydeez 2y agoCUDA is a software moat. If you want to use any gpu other than nvidia, you need to double your engineering budget because theres no easy to bootstrap projects at any level. The hardware prices are meaninglesz if you need a 200k engineer, if they exist, just.to bootstrap a product.
- rbanffy 2y agoDepending on your hardware budget, the engineering one can look like a rounding error.
- cyanydeez 2y agoSure, but then youre still on the.side.of NVIDIA because you jave the.budget.
- sangnoir 2y agoWhy give any additional money to Nvidia when you can announce more profits (or get more compute if you're a government agency) by hiring more engineers to enable AMD hardware for less than a few million per year? It's not like Microsoft loves the idea of handing over money to Nvidia if there is a cheaper alternative that can make $MSFT go up.
- sliken 2y agoSay your success rate for replicating CUDA+Nvidia hardware on AMD is 60%. But it will take 2 years. That's not going to be compelling for any large org, especially when the MI300x is cheaper, but not crazy cheaper than an h100. Especially since CUDA is still rolling out new functionality and optimizations, so the goal posts will keep moving.
- rbanffy 2y agoIt depends on how cheaper the total solution is and how available the hardware is. If I can't get Nvidia hardware less than six months after I get AMD hardware, I have a couple months to port my software to AMD and still beat my competitor that's waiting for Nvidia. It's always a matter of how many problems can you solve for a given amount of money x time.
- cyanydeez 2y agoSure, but the "it depends" is carrying a lot of weight. NVIDIA's moat will get you testable software straight out the gate; any other stack currently is a game of "how long can we take to get this going". Corporations simply arn't interested in long term gains unless there's a straightforward path.
- rbanffy 2y agoIt depends on the problems you have. If you need CUDA, then you married yourself to Nvidia. If you can use libraries that work equally well on both, then you would benefit. When you are a government agency, it’s more palatable to spend the budget in a way it results in employment of nationals and development of indigenous technologies.
- sangnoir 2y ago> Say your success rate for replicating CUDA+Nvidia[...] Rational hyperscalers would just stop as soon as their tooling/workloads/models are functional on AMD hardware within an acceptable perf envelope - just like they already do with their custom silicon. Replicating CUDA is just unnecessary, expensive and time-consuming completionism; if some workloads require CUDA, they will be executed on Nvidia clusters that are part of the fleet.
- cyanydeez 2y agoBecause if you don't join NVIDA your likely hood of success goes down. So the "more profits" you speak of is gambling money. Most corporations arn't going to gamble.
- rbanffy 2y agoDepends on you needing CUDA or not. If you don’t, you can use anything. It was this same game with x86 and ARM is eroding the former king’s place in the datacenter.
- noelwelsh 2y agoLeadership lacking vision + being almost bankrupt until relatively recently.
- hedgehog 2y agoAs another commenter points out their strategy appears to be to focus on HPC clients where AMD can focus providing after-sale software support around a relatively small number of customer requests. This gets them some sales while avoiding the level of organizational investment necessary to build a software platform that can support NVIDIA-style broad compatibility and good out-of-the-box experience.
- dagw 2y agoCUDA is very far ahead. Not only technically, but in mindshare. Developers trust CUDA and know that investing in CUDA is a future proof investment. AMD has had so many API changes over the years, that no one trusts them any more. If you go all in on AMD, you might have to re-write all your code in 3-5 years. AMD can promise that this won't happen, but it's happened so many times already that no one really believes them. Another problem is simply that hiring (and keeping) top talent is really really hard. If you're smart enough to be a lead developer of AMDs core Machine Learning libraries, you can probably get hired at any number of other places, so why choose AMD. I think the leadership gets it and understand the importance, I just don't think they (or really anybody) knows how to come up with a good plan to turn things around quickly. They're going to have to commit to at least a 5 year plan and lose money each of those 5 years, and I'm not sure they can or even want to fight that battle.
- martinpw 2y ago> Another problem is simply that hiring (and keeping) top talent is really really hard. Absolutely. And when your mandate for this top talent is going to be "go and build something that basically copies what those other guys have already built", it is even harder to attract them, when they can go any place they like and work on something new. > I think the leadership gets it and understand the importance, I just don't think they (or really anybody) knows how to come up with a good plan to turn things around quickly. Yes, it always puzzles me when people think nobody at AMD actually sees the problem. Of course they see it. Turning a large company is incredibly hard. Leadership can give direction, but there is so much baked in momentum, power structures, existing projects and interests, that it is really tough to change things.
- deleted 2y ago[deleted]
- DaoVeles 2y agoCUDA is one area that Nvidia really nailed. When it was first announcement I saw it as something neat but could have never envisioned just how ingrained it would become. This was long before AI training/execution was something really on most people radars. But for years I have heard the same things from so many people working in the field. "We hate Nvidia because they got it so right but are the only option."
- alecco 2y ago> are they lacking the right people in senior leadership roles to see this through? Just like Intel, they have an outdated culture. IMHO they should start a software Skunk Works isolated from the company and have the software guys guide the hardware features. Not the other way around. I wouldn't bet money on either of them doing this. Hopefully some other smaller, modern, and flexible companies can try it.
- deleted 2y ago[deleted]
- pjmlp 2y agoYes, to add to the other comments, what many don't realize is that CUDA is an ecosystem, C, C++ and Fortran foremost, however NVidia quickly realized that supporting any programming language community to target PTX was a very good idea. Their GPUs were re-designed to follow C++ memory model, and many NVidia engineers are seat at ISO C++, yet making CUDA the best way to run heterogenous C++. Something that Intel also realized, by acquiring CodePlay, key players in SYCL, and also employing ISO C++ contributors. Then there are the Visual Studio and Eclipse plugins, and graphical debuggers that allow even to single step shaders if you so wish.