2 ms·
The bottleneck is memory bus bandwidth as current Neural Net scaling is parameter count over compute so most A100s only operate at 60% of capacity. Operator fu
by deepnet 4y ago
The bottleneck is memory bus bandwidth as current Neural Net scaling is parameter count over compute so most A100s only operate at 60% of capacity.
Operator fusion ( function composition ) has helped but devs time is impacted by the ever growing list of fused operations.
This is a great overview of what Pytorch 2.0, TorchInduction, and OpenAI's Triton potentially bring to the table.
Namely compilation to LLVM bypassing CUDA, fusing operators and compiling graphs at a lower level allowing other arcitectures, kept out by CUDA's NVIDIA exclusivity, to potentially offer competeting compilers - other compilation toolchains are very early days though.
Backward comptibility is vital and the article details that these innovations are succesfully tested against existing Neural Net code.
NVIDIA still rules the roost in science and industry compute. In this rapidly growing market competition is healthy so this is a good thing for everyone - especially those differentiators like ARM competing on Watts per FLOP.