3 ms·
There's a lot being said about the moat nVidia has with CUDA, but when companies like Tesla are forking $300M on 10,000 GPUs [1], it can't be that difficult to
by notfried 3y ago
There's a lot being said about the moat nVidia has with CUDA, but when companies like Tesla are forking $300M on 10,000 GPUs [1], it can't be that difficult to port whatever code to whatever platform if 100s of millions are in question, or is it?
Can Google/Tesla/etc create an H100-like GPU in a couple of years at a fraction of the cost? And is CUDA really necessary if this much money is at stake?
[1] https://www.tomshardware.com/news/teslas-dollar300-million-ai-cluster-is-going-live-today https://www.tomshardware.com/news/teslas-dollar300-million-a...
- varelse 3y ago[dead]
- smoldesu 3y ago> it can't be that difficult to port whatever code to whatever platform if 100s of millions are in question, or is it? > in a couple of years at a fraction of the cost? And is CUDA really necessary if this much money is at stake? I mean, that's the thing; is resisting CUDA really worth the cost when this much money is at stake? It's not like you have to pay to license CUDA. The two big drawbacks are that it's proprietary and locks you into their ecosystem; neither of which really matter when shipping stuff at that scale. These companies could spend a few million dollars to accelerate their specific codepath, but it's frankly wasted money unless you have a specific reason to avoid Nvidia. The eventual Nvidia-killer will probably be platform-agnostic tooling like ONNX, unless hardware manufacturers revive a sort of OpenCL-style acceleration library.
- 7speter 3y agoIsnt this what intels trying to do with OneAPI?
- smoldesu 3y agoYeah, but a lot of other companies have similar libraries with varying levels of commitment/coverage. ARM has ARMnn, Apple has CoreML, Intel has both Vino and OneAPI (plus SSE/AVX implementations), Google has NNAPI for Android and Google Cloud tensor accelerators, AMD has... Vitis/Xilinix, ROCm, OpenCL and MIGraphX, Windows has DirectML/Olive alongside Microsoft's Azure Execution Providers, Rockchip has RKNPU and Huawei has CANN. Suffice to say, there are a lot of "competing standards" a-la the XKCD : https://xkcd.com/927/ https://xkcd.com/927/ I don't think any of those will really topple CUDA, though. The best way to unseat Nvidia's dominance would be to target their two weaknesses; the closed nature of CUDA and the lock-in to Nvidia hardware. It would be difficult to overturn them both, but an open and fast CUDA alternative is really all people actually want. That's why I think libraries like ONNX have the right idea; instead of relying on chip manufacturers to not rip each other's throats out (they won't), they unify everyone's proprietary APIs. Barring some ground-up GPGPU library like OpenCL, this seems like the smartest path to me.
- emadm 3y agoThe CUDA moat argument just doesn’t hold up in real life for foundation model training and inference We get equivalent performance on non-NVIDIA chips without it and most of the stuff is abstracted away these days We do write CUDA when needed but really only a handful of folk will actually train models as it is a pain
- smoldesu 3y agoCUDA's moat is not built around inference, as far as I can tell. It's a very handy tool for deployments, and often gives them the edge when Nvidia pulls ahead, but it isn't a be-all-end-all. The inferencing situation has been pasted over by countless vendor-specific libraries and stuff like Pytorch and ONNX. The real reason for the moat comes down to a number of things, like: - The flexibility of generic GPU and ML acceleration primitives - The availability of systems with hundreds of terabytes of GPU memory for you to scale to - The struggle of trying to use commodity hardware for actual ML acceleration I deploy models to freely-provisioned ARM servers, I don't think I'm a choosing beggar in the slightest. When I deploy to Nvidia hardware though, the experience is much nicer on-the-whole. This stuff is definitely possible with consumer hardware, ROCm acceleration or CoreML optimization, no doubt. It's not hard to see why Nvidia is at the mountaintop right now though, and unless the industry agrees to stop building CUDA-style moats then this is the history we're damned to repeat.
- 7speter 3y agoDunno if this is a rhetorical question because a lot of people here seems to think that yes, these companies can make their own silicon for much cheaper. The problem is that they need time, and right now they still have work to do, so the best way to solve this problem is to buy from nvidia. Or am I saying too much?
- kiratp 3y ago> it can't be that difficult to port whatever code to whatever platform if 100s of millions are in question, or is it? 300M hires a lot of engineers. Maybe the fact that they are buying instead of building tells you how hard it is to build and how strong a moat Nvidia has. Musk has publicly said that the true test of their own training hardware efforts is if the ML team switches to it. I.e. it’s a software, not hardware problem.
- coolspot 3y agoGoogle has TPU v5 which has performance comparable to H100 .
- outside1234 3y agoHonest question - why is this not a bigger competitor then?
- dotnet00 3y agoMaybe flexibility? Google's TPUs are comparable in performance for pure ML, but at the same time if you're building a huge expensive cluster, you might want to be able to have the flexibility to accelerate other tasks (eg scientific compute).
- blovescoffee 3y agoBecause everything effectively trickles down from what academics/researchers are doing and they're not using google TPUs in the same quantities.
- shaklee3 3y agoYou can't buy TPUs and put them in your own data center
- coffeebeqn 3y agoI’d be surprised if we don’t see more specialized chips in the future- kind of like Bitcoin isn’t mined on GPUs anymore. There just hasn’t been much of a need for specialized AI chips so far but GPUs have been useful and profitable in rendering etc for a long time. But I think Google made “TPU”s already (tensor processing unit)