3 ms·
All the reasons: [1] The compilers don’t produce great instructions; [2] The drivers crash frequently: ML workloads feel experimental; [3] Software adoption
by ternaus 3y ago
All the reasons:
[1] The compilers don’t produce great instructions;
[2] The drivers crash frequently: ML workloads feel experimental;
[3] Software adoption is getting there, but kernels are less optimized within frameworks, in particular because of the fracture between ROCm and CUDA. When you are a developer and you need to write code twice, one version won’t be as good, and it is the one with less adoption;
[4] StackOverflow mindshare is lesser. Debugging problems is thus harder, as fewer people have encountered them.
---
were crucial when we had enough supply of NVidia GPUs, but if demand described in https://gpus.llm-utils.org/nvidia-h100-gpus-supply-and-demand/ https://gpus.llm-utils.org/nvidia-h100-gpus-supply-and-deman...
is real (450,000+ H100)
Software bottlenecks most likely will be addressed sometime soon