23 ms·
The CUDA moat argument just doesn’t hold up in real life for foundation model training and inference We get equivalent performance on non-NVIDIA chips without
by emadm 3y ago
The CUDA moat argument just doesn’t hold up in real life for foundation model training and inference
We get equivalent performance on non-NVIDIA chips without it and most of the stuff is abstracted away these days
We do write CUDA when needed but really only a handful of folk will actually train models as it is a pain
- smoldesu 3y agoCUDA's moat is not built around inference, as far as I can tell. It's a very handy tool for deployments, and often gives them the edge when Nvidia pulls ahead, but it isn't a be-all-end-all. The inferencing situation has been pasted over by countless vendor-specific libraries and stuff like Pytorch and ONNX. The real reason for the moat comes down to a number of things, like: - The flexibility of generic GPU and ML acceleration primitives - The availability of systems with hundreds of terabytes of GPU memory for you to scale to - The struggle of trying to use commodity hardware for actual ML acceleration I deploy models to freely-provisioned ARM servers, I don't think I'm a choosing beggar in the slightest. When I deploy to Nvidia hardware though, the experience is much nicer on-the-whole. This stuff is definitely possible with consumer hardware, ROCm acceleration or CoreML optimization, no doubt. It's not hard to see why Nvidia is at the mountaintop right now though, and unless the industry agrees to stop building CUDA-style moats then this is the history we're damned to repeat.