3 ms·
> The big question is whether demand for CUDA can be supplanted with application-specific accelerators. At least for AI workloads, Google's XLA compiler and th
by felarof 2y ago
> The big question is whether demand for CUDA can be supplanted with application-specific accelerators.
At least for AI workloads, Google's XLA compiler and the JAX ML framework have reduced the need for something like CUDA.
There are two main ways to train ML models today:
1) Kernel-heavy approach: This is where frameworks like PyTorch are used, and developers write custom kernels (using Triton or CUDA) to speed up certain ops.
2) Compiler-heavy approach: This uses tools like XLA, which apply techniques like op fusion and compiler optimizations to automatically generate fast, low-level code.
NVIDIA's CUDA is a major strength in the first approach. But if the second approach gains more traction, NVIDIA’s advantage might not be as important.
And I think the second approach has a strong chance of succeeding, given that two massive companies—Google (TPUs) and Amazon (Trainium)—are heavily investing in it.
(PS: I'm also bit biased towards approach 2), we build llama3 fine-tuning on TPU https://github.com/felafax/felafax https://github.com/felafax/felafax)
- 01100011 2y agoIt's weird to me that folks think NVDA is just sitting there, waiting for everyone to take their lunch. Yes, I'm totally sure NVDA is completely blind to competition and has chosen to sit on their cash rather than develop alternatives...</s>
- ein0p 2y agoNot really, no. Over the past several years, JAX was used in only 3% of top publications. PyTorch in 60%. There's no trend to suggest that JAX has "reduced the need" for anything, except for Google itself. https://paperswithcode.com/trends https://paperswithcode.com/trends