11 ms·
CUDA Is Still a Giant Moat for Nvidia
- mrbishalsaha 3y agoWell deserved in my mind. Nvidia has been pushing the use of AI chips for far to long. The literally did everything possible to make it happen. I believe LLMs will be commoditised while the compute power will be the next big thing.
- aurareturn 3y ago>I believe LLMs will be commoditised while the compute power will be the next big thing. Can you talk more about this? Would love to understand.
- deleted 3y ago[deleted]
- chii 3y ago> Well deserved in my mind. not if this moat could be leveraged into a monopoly on AI chips, to the detriment of society. I want to see competition in this space. Unfortunately, the market rally of nvidia stock is suggesting that most investors are expecting this monopoly to eventuate. Therefore, it is in the interest of society to ensure that such a software moat is not established. Look what happened to the web browser when microsoft held a monopoly on it, and look at what is happening with chrome, apple appstore, etc.
- anon291 3y ago> Look what happened to the web browser when microsoft held a monopoly on it, and look at what is happening with chrome, apple appstore, etc. Realistically what happened is that after a few decades of development, competitors arose and took the market. In the meantime, Microsoft became rich. Who cares
- andsoitis 3y agoif the prize is big enough, there will rise others.
- chii 3y agoso why isn't there a windows competitor today for the consumer PC market? Surely it's a bigger market than ai chips. The answer is that the entrenchment of the tools, software and inertia of a defacto standard is what prevents new entrants. The time to stop it is to nip it in the bud. Prevent monopoly from forming, rather than hope that after the monopoly forms, some competitor will break it.
- dcgudeman 3y agoWhat do you suggest? Kneecap the most competent company executing in AI hardware to help laggards compete?
- chii 3y agonobody is kneecapping anyone. The ask is to change CUDA into a standard for which nvidia is one implementation. instead, what you have today is this: > You may not reverse engineer, decompile or disassemble any portion of the output generated using SDK elements for the purpose of translating such output artifacts to target a non-NVIDIA platform. from https://docs.nvidia.com/cuda/eula/index.html#limitations https://docs.nvidia.com/cuda/eula/index.html#limitations section 8
- bsder 3y agoCUDA is a moat because AMD and Intel are run by morons^W^W^W run by people who can't swallow the fact that software is more important than hardware. Intel should be shoveling out 16GB Arc graphics cards for free to every graduate program in the country who can fill out a web form. In a couple years, they'd displace NVIDIA. AMD needs to be funding a CUDA shim that allows people to port stuff directly to their cards. And they need to NOT be segmenting the consumer and professional cards software ecosystems. Yes, there has been progress. However, when you look at the amount of money that AMD and Intel throw at software vs how much NVIDIA throws at software, it's an instant facepalm moment. NVIDIA is 100% vulnerable--if it weren't for the fact that their competitors are idiots.
- froonly 3y agoCUDA is a shallow moat whose effectiveness depends entirely on NVidia convincing people to be mortally fearful of water.
- kkielhofner 3y agoGenuine questions. What are your use cases? What do you do? How much experience? My personal experience shows CUDA to in fact be a very deep moat. In ~12 years CUDA and ~6 ROCm (since Vega) I’ve never met a professional who says otherwise, including those at top500.org AMD sites. From what I’ve seen online this take really seems to come from some kind of Linux desktop Nvidia grudge/bad experience or just good ‘ol gaming/desktop team red vs green vs blue nonsense. Many things can be said about Nvidia and all kinds of things can be debated but suggesting that Nvidia has > 90% market share simply and solely because people drink Nvidia kool-aid is a wild take.
- froonly 3y agoI have 40+ yrs of HPC/AI apps/performance engineering experience & I was one of the 1st people to port LAPACK and a number of other numerical libs to CUDA. Moreover, many of those major DoE + AI sites are my customers. You should not confuse AMD's general & long-standing indifference/incompetence wrt SW with the actual difficulty of providing a portable SW path for acceleration. As Woody Allen once said: "90% of success is showing up" But what happened in AI, when, in a very short period of time, almost everyone moved away from writing their directly in CUDA, to writing them in frameworks like Tensorflow & PyTorch is all the evidence anyone need to show just how unsound that SW obstacle is.
- frozenport 3y ago+1 on the AMD are morons train. AMD recently got rid of one of the CUDA compatibility layers instead of extending it.
- pests 3y agoThey didn't get rid of it, they dropped development and released it as open source.
- steelbrain 3y ago> they dropped development and released it as open source. "They" (being AMD) didn't. The person they contracted put in a clause that allowed him to open source the work (years AFTER) AMD stopped paying him.
- pests 3y agoI don't think the person put in the cluse by himself, surely AMD would have known and agreed to it. They would own the output of any of his work anyways, it was only with their approval was he was allowed to take ownership or release it. I do think this is splitting hairs, AMD isn't the savor here but I wouldn't say they don't get any credit either.
- modeless 3y agoChasing compatibility is a waste of time and ultimately counterproductive. The important software is open source, they can just add direct support for their stuff. What they need to do is fix the stability of their drivers, make their stuff work on every GPU they sell or have sold in the past few years (as CUDA always has), and pay employees to integrate support into all the popular open source projects while fixing every bug that gets reported. And they need to release high-RAM versions of their next gaming GPUs. More than anything else that will incentivize people to switch. If they're selling 36 GB while Nvidia is still selling 24 GB, people will do what it takes to move over.
- jjmarr 3y ago
- fancyfredbot 3y agoI am not sure CUDA is the moat, but yes, software is the moat. To first order nobody writes any CUDA, and even if you do you are probably bad at it. The language is slightly easier to use than openCL but writing really performant code is still a nightmare (a pipeline of asynchronous memory copies from global to shared memory is not easy to program but this is a requirement for full performance on tensor cores). So no, the moat really isn't the language. It's not even the libraries, it's the integration of the libraries into third party software like pytorch, jax, etc. This is the truly massive advantage NVIDIA has, and they got it by being early and by being installed in an awful lot of machines.
- frozenport 3y agoYeah in the ML space you don't need to, but in engineering, HPC its still really popular. Perhaps in some universe we'll replace C++ with ONNX.
- homarp 3y agoor use tvm
- danielscrubs 3y ago”To first order nobody writes any CUDA and even if you do you are probably bad at it” is such an anti-intellectual stance that is repeated to such a large extent that it irks me. It’s the authors protecting their ego and is said about everything they don’t understand. It is said about compilers, about static typing, about pretty anything the authors do not yet know. At least say why people wouldn’t be good at it. The documentation is poor, the GPUs are a black box or anything in that vein. Then they can help you learn instead of preemptively dismiss it.
- aurareturn 3y agoSo do people write in CUDA? I assume non-ML scientists do but ML researchers don't?
- jgord 3y agoI dont understand this - arent almost all ML NN models built in pytorch, and arent these compiled / jit'd into a lower level format - and can we not have various backends/drivers for that, such as CUDA / ROCM / vnni ? The article is unsatisfying because it doesnt explain WHY cuda reigns supreme. One hypothesis put forward is that the main alternative ROCM is just not very complete and not very fast - thats a good argument. Another hypothesis that is not considered is : CUDA reigns supreme, because NVIDIA GPUs reign supreme. But people dont write CUDA code .. they write pytorch code ?!
- dannyw 3y agoNobody else seems to be willing to invest serious funding, including market rates for SWEs, into compelling alternatives. I believe AMD's TC for senior software engineers tops out at 200k in the Bay Area. The problems you generally experience are: * Inexplicably poor performance * Poor (and sometimes incorrect) documentation * Difficulties debugging * Crashes and hangs
- anon291 3y agoI'm an AI compiler engineer and AMDs hiring process was ... Non-competitive. Companies are hiring left and right at a fast clip and heres AMD wanting you to fly out in a month. I love their CPUs but... Come on. You gotta be serious to compete
- KerrAvon 3y agoYou really mean TC and not base salary? That’s shockingly bad.
- kimixa 3y agoIt's also not correct. I wouldn't consider myself a "Senior" engineer, but am at AMD and have a TC notably higher than that.
- aurareturn 3y agoWhy is this? AI is going to a multi trillion market. I can't think of anything else bigger except maybe electricity, real estate, and food. If I'm AMD, I'd spend at least $1 billion/year figuring out the software side. I can't think of an easier way for AMD to return value to shareholders than eroding CUDA advantage. Heck, Meta invested something like $100b on VR so far and VR is not nearly the market that AI is.
- aurareturn 3y agoChip War has a great section on how the Soviet Union tried a “just copy/steal” strategy in semiconductors and fell hopelessly behind because of it. It’s a great theoretical idea to just copy/steal and fast-follow, but semiconductors, AI, and other “harder technologies” require building human and intellectual capital that will get better with time. From there, you need to have the prior generation to keep up with ever-increasing complexity and difficulty as these things get more advanced. I disagree with your section on Huawei and China. China isn't just trying to just copy/steal AI. In terms of models, China is a bit behind in LLMs but arguably more ahead in self-driving cars. China is throwing everything at semiconductor manufacturing instead because that's where their bottleneck truly is - not CUDA. Had Huawei had access to TSMC's 5nm and 3nm, they might already be equal to Nvidia in raw GPU prowess. After all, HiSilicon's Kirin already matched/exceeded Qualcomm before the Trump ban. Their 5G chips/implementation were well ahead of anyone else. In software, it's easier for China to adopt a CUDA alternative because China is usually really good at unifying under one vision - especially when they have to.
- rikafurude21 3y agoAnyone following geohots current tinygrad struggles has seen this proven right in front of them. AMD gpus are practically unusable for any serious ML work and he had to learn it the hard way, having dropped 100K into AMD gpus, assuming the drivers would work and if not, he even personally offered to fix them.
- Buttons840 3y agoI saw him on Twitch today in passing, the title was about "ripping <something> out of AMD drivers" or similar, so it seems he's still at it.
- KennyBlanken 3y agoIs this the same geohot that 9+ months ago declared he was "done with AMD"? Isn't geohot infamous for stealing other people's work? PBCAK? That said, ROCm only officially supports a fraction of its product line, and an odd smattering throughout at that. It's a joke compared to CUDA which will run on damn near anything. And AMD has a long, long history of dogshit drivers (at least on Windows.) AMD just doesn't seem to give enough of a shit to invest money into securing top talent for this, and NVIDIA will continue to stomp them.
- xiphias2 3y agoThat same bug is still open, not fixed. Azure announced access to AMD GPU cloud with NDA, but the cards are unusable for compute work as they lock up randomly.
- justinclift 3y ago> Isn't geohot infamous for stealing other people's work? Are you meaning the Sony Playstation hacking where they took legal action against him, or are you meaning other stuff?
- arcanemachiner 3y agoIn lieu of a real answer, here's more HN conjecture: https://news.ycombinator.com/item?id=30740509 https://news.ycombinator.com/item?id=30740509
- shmerl 3y agoLock-in should be broken. CUDA is one of the worst things about this whole ecosystem. Looks like AMD came close to breaking it, but they abandoned developing the translation layer.
- anon291 3y agoThinking that a cuda translation layer will take away Nvidias advantage is like expecting writing a c compiler to spontaneously result in unix
- shmerl 3y agoIt would take away a huge chunk of their advantage, no doubt about it. Let Nvida compete on merit instead of lock-in. Then you can say their advantage lies in being better. But Nvidia is very lock-in oriented, which undermines the claim that they are so much better than everyone.
- buildartefact 3y agoCUDA IS the merit. They’ve been developing the software stack to make GPU programming accessible for 2 decades now. They’re a software company as much or more than they are a hardware company. Ignoring this fundamentally misunderstands why they’re in the position they’re in now.
- shmerl 3y agoIt's not the merit - it's the moat (i.e. lock-in) as the linked article states. Merit in this context would mean something you can compare across different GPUs. For CUDA - you can't. I.e. it's a tool to force you to use Nvidia.
- sbierwagen 3y agoIf I was an AMD shareholder I'd seriously be considering a vote to remove CEO Lisa Su. They make nearly identical products to NVIDIA, yet that other company is worth literally ten times as much, because pytorch actually works on their cards. Why isn't she prioritizing firmware that doesn't crash?
- david-gpu 3y ago> Why isn't she prioritizing firmware that doesn't crash? I used to work in the GPU industry and this sort of view is both pervasive and misguided. GPUs are immensely complex machines. It is really hard to get them to work, let alone work with high performance. Because of this, and in spite of the amount of time and resources spent on validation and verification, the hardware often contains flaws. It is the responsibility of the drivers to work around these flaws in various ways. When a flaw hasn't been discovered and worked around yet, you perceive it as the GPU being unstable or crashing. There is no fast simple solution to this. You need a finely tuned corporate machine from beginning to end. Better hiring processes, better management, better design processes, better verification processes, better software development practices, better marketing and sales, better customer relations. Everything.
- imtringued 3y ago>GPUs are immensely complex machines. It is really hard to get them to work, let alone work with high performance. This is like saying combustion engines are immensely complex machines when your car suddenly loses power on the highway for no apparent reason and then when you restart the engine it works for another five minutes again. When you drive on normal roads it works flawlessly. It must be the engine, right? After all, it is the most complicated aspect! Except in reality it is far more likely for it to be a problem in the electronics driving the fuel pump or spark plug. AMD most likely has some sort of buffer overflow or deadlock in their GPU drivers that is causing difficult to diagnose problems. It is very unlikely that the silicon itself is broken when it works fine for playing video games and it also works fine when your GPU is one of the few officially supported by ROCm.
- 3y ago
- TMWNN 3y agoHow does Apple's Metal compare to/compete with CUDA? I know Ollama and LM Studio support Metal.
- Havoc 3y agoBlows my mind that AMD isnt throwing everything they’ve got at fixing this.
- Alifatisk 3y agoBy reading all the comments here, everyone seem to agree on that AMD is betting on the wrong thing. Yet, they continue the same path. - Abandoning ZLUDA was maybe not the best choice - Not accepting the fact that software is equally as important as hardware is wrong - Pushing more vram into their cards would attract more people - Fix hardware issues (especially with the restarts on every fail) should be high priority
- beryilma 3y agoI don't know much about CUDA and NVIDIA, but it has always surprised me how hardware companies are so bad at producing good software tooling for their hardware. Many microcontroller companies have terrible software support: no free C/C++ compilers, clunky IDEs, too much reliance on 3rd party software providers, no decent code libraries... Even if they have software support, the code is bad and bloated. Look at ST's HAL libraries, for example. Thankfully, an open source or free tool often comes to the rescue, usually through the efforts of dedicated individual programmers. But billion-dollar companies relying on such 3rd party tooling seems insane to me.