5 ms·
Nvidia's biggest advantage in AI has never been only their hardware performance but how entrenched their software is in ML research that flowed down stream. How
by YuechenLi 2mo ago
Nvidia's biggest advantage in AI has never been only their hardware performance but how entrenched their software is in ML research that flowed down stream. However, if you've actually used CUDA C/C++, it's pretty one of the worst software development ecosystem imaginable: you get all the footgun of regular C++, plus GPU compute pretending to be C++ and but doesn't actually behave like C++ because CPU and GPU compute are fundamentally different, and the only reason people put up with it is because Vulkan and HIP C/C++ are even worse.
Google's limitation is that they still don't offer TPUs in a PCI-E card/dev board that people can plug in to their PC for local development and sane low level API to develop against, instead you have to go through their cloud and their full software stack which greatly limits ecosystem growth. The minute that Google figures that out, that's when Nvidia's dominance would be challenged.
- HeWhoLurksLate 2mo agoI mean they had/have the Coral but that's in an entirely different market segment
- bigyabai 2mo agoCoral and the TPUs are ASICs, and therefore are barely reprogrammable. It doesn't really compare to the complexity and flexibility of CUDA ALUs.
- musebox35 2mo agoThe biggest advantage of tpus is the high bandwidth fiber optic interconnect between them that allows distributed computing on pods with thousands of tpus and the co-design of cooling systems that go with their racks. I do not think that we will see personal tpus any time soon.
- ravenstine 2mo agoThat's really interesting. I have no experience writing anything that involves GPUs/TPUs, but over the years I've consistently read that CUDA is the "real moat" of Nvidia, which I never totally believed, but the way you describe makes it seem like it's not actually a moat in the slightest. It just happens to be an ecosystem associated with hardware that is not only considered the gold standard but happens to be more open than potential competition. Could it be that Nvidia has been on top because none of the competition has actually tried kicking them where it hurts?
- csomar 2mo agoSoftware has always been the moat but for some reason it's always hamstrung by upper management. The latest of the frenzies being replacing sane (or whatever we have) of development practices with AI-slop. Management likes it because it removes software developers from the loop.
- szundi 2mo ago[dead]
- calebkaiser 2mo agoIt's a totally reasonable question, and one that everyone asks when they're learning about CUDA. The frustrating answer to your last question is that lots of companies have shipped GPU dev environments that can theoretically be used instead of CUDA. AMD has ROCm, Apple has had a couple projects (OpenCL, Metal), Intel has some stuff, and there are newer efforts like TinyGrad + a generation of slightly higher level frameworks from AI companies, like Triton from OpenAI. The basic problem is that CUDA has become something of a Schelling point. If you want to train a model right now, the highest performance you can get is almost certainly on CUDA. From the basic general matrix multiply operation, to specific NN architectures, CUDA is going to have incredibly optimized implementations out of the box. And it's going to make multi-GPU training so much easier. And all the dependencies you build on (those layers you import from PyTorch or Transformers or whatever) are going to work optimally right away on CUDA. And that weird random repo that you found with a unique optimizer--it runs on CUDA too. And now the cool new implementation that you're about to release is also going to be built for CUDA. It's so tempting to think "Just write replacement software", but you also need to transition the entire ecosystem in large part to match CUDA's effectiveness, and you need to get comparable performance out of your chip/library combo as NVIDIA can get out of its cards with CUDA. There's a whole story here to how effective NVIDIA has been at navigating this. Very early on, they heavily prioritized PyTorch and TensorFlow, getting involved in the projects as much as they could and making sure they always ran best on CUDA. But the TLDR is that yes, you're right, another company could write a CUDA competitor. But actually replacing CUDA is a much larger task. I'm personally hopeful that with the rise of coding agents, we see more movement on this front with other projects moving into view. It will take some time for any ecosystem to start to emerge that can dislodge CUDA for researchers who don't want to dive that deep into the stack, but hopefully we start to see some momentum build.
- whatever1 2mo agoNow with LLMs why a programming framework is a moat?
- tomaskafka 2mo agoI had a hard time understanding why didn’t AMD make a better developer experience for this two years ago, and am now even more baffled that even with all the LLMs they still don’t seem to have moved a single inch, despite this probably being a tens of billions dollars worth feature.
- kllrnohj 2mo agoLLMs make the dev environment almost irrelevant. llama.cpp supports AMD with both ROCm and Vulkan and that's nearly all that matters now. TBD how much AMD's AI Halo play will change things if at all, but they got a lot of positive press in launch reviews for having an actually robust software story for once.
- CorrectHorseBat 2mo agoHardware companies are notoriously bad at software. The software they use sucks, the languages they use suck, the internal tooling software they write sucks. They don't know what good developer experience is, how do you expect them to deliver it to other people?
- tomaskafka 2mo agoI understand this; I would expect that a $50bn carrot would suffice to move the donkey, but it seems it does not.
- npunt 2mo agoAre the switching costs of CUDA ecosystem potentially threatened because LLMs are now quite good at transcoding into other languages? In other words, is Nvidia's greatest strength (AI) also potentially its undoing?
- schopra909 2mo agoI’m not entirely sure if local development will lead to Nvidia’s supremacy being challenged. I think a simple reason why it’s been hard to unseat in Nvidia is first mover advantage. A lot more water has flown through Nvidia pipes than TPUs or AMDs chips for that matter. TPUs and AMD chips aren’t priced cheaper than NVIDIA (at least for my purposes training models). So there hasn’t been an impetus for me to venture there and use those chips. Anecdotally, folks I know who have tried using TPUs and AMD chips have hit more issues with the underlying drivers than with NVIDIA chips. That costs time and money to fix. Eventually the other chips will go through enough iterations and stability will be reached
- ijidak 2mo agoGenuine question. Given that LLMs are supposed to allow us to rewrite anything, and I am an LLM believer, what I don't understand is: how does CUDA continue to be a moat in a world where LLMs can rewrite entire software development stacks? If NVIDIA is right about AI, isn't this same technology going to erode the software side of this same software moat?
- ceehex 2mo agowell done you realised no one knows what they are talking about
- gr_norm 2mo agoIndeed. Now, given that CUDA is apparently not being usurped, update your priors.
- polanyer 2mo agoIt seems like path dependent lock in to me and a risk/reward calculation. What do you gain by not using CUDA vs what do you risk?
- robocat 2mo ago> Google's limitation is that they still don't offer TPUs in a PCI-E card/dev board Nvidia's sells hardware yet their market cap is about the same as Google's. How much value could Google get by selling hardware too? Google'd be selling to competitors, so difficult to capture much of the value and would decrease Google's value as an AI company. Maybe a child company?
- bdangubic 2mo agoGoogle sells TPUs already, to competitors :)
- akoboldfrying 2mo agoInteresting take on Google's TPUs. What I've previously heard (and still believe) is that Google's decision to only rent out, never sell, their TPUs is a deliberate and savvy strategy for bolstering GCP, which will work provided that TPUs are able to actually compete with other hardware (in practice meaning Nvidia). A few months ago there was some discussion on HN comparing them, and I think the verdict at the time was that their latest-gen TPUs win on compute-per-Joule for LLM-type workloads by quite a margin, which I think is huge for those who want to run LLMs at scale.
- bri3d 2mo agoThe CUDA runtime coming with a gazillion reasonably decent kernels (DNN, BLAS, CUTLASS) and a concurrency system (NCCL) is a big deal; especially in the “early days” very few researchers or development runtimes were even writing their own kernels or dealing with CUDA C++ extensively, they were wrapping the ones NVidia gave them. I do agree that it’s really not great, and I also have never been a strong believer in the CUDA moat overall; as the need for GPUs moves from research to production (inference), companies are plenty willing to build software from scratch anyway (and we see this with AMD GPUs being in plenty high demand in the datacenter and enthusiast market now).
- galaxyLogic 2mo agoWhat I don't quite get is why can't they use AI to translate CUDA programs into more open architectures like AMD ROCm? AI is supposed have solved the "coding problem". But shouldn't translating a program from one platform to another be an even easier, more mechanical, task for the AI?
- mdp2021 2mo ago> AI is supposed have solved the Which AI? LLMs are coding facilitators and code producers. A problem is solved when the solution is reliable. Non-deterministic Neural Networks are not reliable. In fact, > more mechanical[] task that suggests an expectation of process and procedure, which is still not a capability of current architectures. Sure, you can ask a brains-deficient operator to perform a huge task, but then you'll have to check the whole product, and that remains not cheap.
- larnon 2mo agoThe tool itself (AI) may not be reliable, but that is also very true for every other tool (e.g. Human). Also, you are right about when a problem is solved, but this doesn't need the tool to be reliable as you said, just the solution part. Hence, as long as the produced code works as intended, it doesn't matter what you used to produce the output.
- Someone 2mo ago> Google's limitation is that they still don't offer TPUs in a PCI-E card/dev board that people can plug in to their PC for local development I’m not familiar with this field, but to my brain, https://www.amazon.com/s?k=Google+Coral https://www.amazon.com/s?k=Google+Coral seem to show me several such options.
- spwa4 2mo ago... with performance greater or comparable to even the lowest performance Nvidia card?