5 ms·
The way I see, NVidia only has a few advantages ordered from most important to least: 1. Reserved fab space. 2. Highly integrated software. 3. Hardware archi
by Laremere 3y ago
The way I see, NVidia only has a few advantages ordered from most important to least:
1. Reserved fab space.
2. Highly integrated software.
3. Hardware architecture that exists today.
4. Customer relationships.
but all of these aspects are weak in one way or another:
For #1, fab space is tight, and NVidia can strangle its consumer GPU market if it means selling more AI chips at a higher price. This advantage is gone if a competitor makes big bets years in advance, or another company that has a lot of fab space (intel?) is willing to change priorities.
2. Life is good when your proprietary software is the industry standard. Whether this actually matters will depend on the use case heavily.
3. A benefit now, but not for long. It's my estimation that the hardware design for TPUs is fundamentally much simpler than for GPUs. No need for raytracing, texture samplers, or rasterization. Mostly just needs lots of matrix multiplication and memory. Others moving into the space will be able to catch up quickly.
4. Useful to stay in the conversation, but in a field hungry for any advantage, the hardware vendor with the highest FLOPS (or equivalent) per dollar is going to win enough customers to saturate their manufacturing ability.
So overall, I give them a few years, and then the competition is going to be real quite fast.
- 7e 3y agoThese have always been NVIDIA's "few" advantages and yet they've still dominated for years. It's their relentless pace of innovation that is their advantage. They resemble Intel of old, and despite Intel's same "few" advantages, Intel is still dominant in the PC space (even with recent missteps).
- weweersdfsd 3y agoThey've dominated for years, but now all big tech companies are using their products in scale not seen before, and all have vested interest in cutting their margins by introducing some real competition. Nvidia will do good in the future, but perhaps not good enough to justify their stock price.
- logicchains 3y ago>2. Highly integrated software. NVidia's biggest advantage is that AMD is unwilling to pay for top notch software engineers (and unwilling to pay the corresponding increase in hardware engineer salaries this would entail). If you check online you'll see NVidia pays both hardware and software engineers significantly more than AMD does. This is a cultural/management problem, which AMD's unlikely to overcome in the near-term future. Apple so far seems like the only other hardware company that doesn't underpay its engineers, but Apple's unlikely to release a discrete/stand-alone GPU any time soon.
- nl 3y agoActually their real advantage is the large set of highly optimised CUDA kernels. This is the thing that lets them outperform AMD chips even on inferior hardware. And the fact that anything new gets written for CUDA first. There is OpenAI's Triton language for this too and people are beginning to use it (shout out to Unsloth here!). > Reserved fab space. While this is true, it's worth noting that the inference only Groq chip which gets 2x-5x better LLM inference performance is on a 12nm process.
- ants_everywhere 3y agoHonest question: will AI help AMD catch up with optimized CUDA/ROCM kernels of their own?
- dagmx 3y agoDon’t underestimate CUDA as the moat. It’s been a decade of sheer dominance with multiple attempts to loosen its grip that haven’t been super fruitful. I’ll also add that their second moat is Mellanox. They have state of the art interconnect and networking that puts them ahead of the competition that are currently focusing just on the single unit.
- latchkey 3y agoThis moat is going to get paralleled over the next few years. First off Mellanox is unobtanium with 52+ week lead times. GigaIO has a PCIe fabric solution that is a fraction of the cost of Mellanox and available today. This enables up to 64 GPUs to appear on a single system. We're also seeing the ultraethernet stuff come online as well, but that'll have to wait for PCIe6.
- otabdeveloper4 3y agoCUDA is absolute shit, segfaults or compiler errors if you look at it wrong. NVidia's software is the only reason I'm not using GPU's for ML tasks and likely never will.
- KeplerBoy 3y agoThat's just C. If you're accessing your arrays out of bounds it's going to segfault. hopefully. Can't blame CUDA for that one.
- otabdeveloper4 3y agoI'm talking about the compiler segfaulting, not the end-user code.
- Culonavirus 3y agoSkill issue.
- otabdeveloper4 3y agoNo, CUDA's botched gcc implementation segfaulting due to compiler errors during compilation is not a "skill issue". (Well, a skill issue of whoever is patching gcc on Nvidia's end, I guess.)
- KeplerBoy 3y agoNvidia's datacenter AI chips don't have raytracing or rasterization. Heck, for all we know the new blackwell chip is almost exclusively tensor cores. They gave no numbers for regular CUDA perf.
- deleted 3y ago[deleted]
- bartwr 3y agoSeems you have not worked with ML workloads, but base your comment on "internet wisdom", or worse, business analysts (I am sorry if that's inaccurate). On GPUs, ML "just works" (inference and training) and are always order of magnitude faster than whatever CPU you have. TPUs work very well for some model architectures (old ones that they were optimized and designed for) and on some novel others can be actually slower than a CPU (because of gathers and similar) - this was my experience working on ML stuff as an ML Researcher at Google till 2022, maybe it got better but I doubt. Older TPUs were ok only for inference of those specific models and useless for training. And anything new I tried (fundamental part of research...) - the compiler would sonetimes just break with an internal error, most of the time just produce terrible and slow code, and bugs filed against it would stay open for years. GPU is so much more than a matrix multiplier - it's a fully general, programmable processor. With excellent compilers, but most importantly - low level access that you don't need to rely on proprietary compiler engineers (like TPU ones) and anyone can develop something like Flash Attention. And as a side note: while a Transformer might be mostly matrix multiplication, many other models are not.
- sevagh 3y agoAlso, it's disingenuous to say "there's only 4 things you need to beat NVIDIA" when each of the 4 is an enormous undertaking.
- puppymaster 3y agonot to mention every not-so-serious, inference heavy ML developers just want something to work to deliver to client. That itself is a semi-moat.
- kkielhofner 3y agoIt's been talked to death but non-CUDA implementations have their challenges regardless of use case. That's what first-mover advantage and > 15 years of investment by Nvidia in their overall ecosystem will do for you. But support for production serving of inference workloads outside of CUDA is universally dismal. This is where I spend most of my time and compared to CUDA anything else is non-existent or a non-starter unless you're all-in on packaged API driven Google/Amazon/etc tooling utilizing their TPUs (or whatever). The most significant vendor/cloud lock-in I think I've ever seen. Efficient and high-scale serving of inference workloads is THE thing you need to do to serve customers and actually have a chance at ever making any money. It's shocking to me that Nvidia/CUDA has a complete stranglehold on this obvious use case.
- Oioioioiio 3y agoNvidia has so much software behind all of this, your list is a tremendes understatement. Alone how many internal ML things nvidia builds helps them tremendesly to understand the market (what does the market need). And they use their inventions themselves. 'only has a few' = 'has a handful easy to list but with huge implications which are not easily matched by amd or intel right now'
- jimberlage 3y agoI’ve spent the last month deep in GPU driver/compiler world and - AMD or Apple (Metal) or someone (I haven’t tried Intel’s stuff) just needs to have a single guide to installing a driver and compiler that doesn’t segfault if you look at it wrong, and they would sweep the R&D mindshare. It is insane how bad CUDA is; it’s even more insane how bad their competitors are.
- jimberlage 3y agoIf you work in hardware and are interested in solving this lemme say this There are billions of dollars waiting for the first person to get this right. The only reason I haven’t jumped on this myself is a lack of familiarity with drivers.