4 ms·
> we used tensor cores and managed to get back fp32 accuracy with 3 rounds of the things Hey are you referring to 3xTF32 (https://github.com/NVIDIA/cutlass/tre
by junrushao1994 4y ago
> we used tensor cores and managed to get back fp32 accuracy with 3 rounds of the things
Hey are you referring to 3xTF32 (https://github.com/NVIDIA/cutlass/tree/master/examples/28_ampere_3xtf32_fast_accurate_tensorop_fprop https://github.com/NVIDIA/cutlass/tree/master/examples/28_am...)? IMO this is a perfect example where proper abstraction could save engineers non-trivial amount of time - imagine a compiler stack which allows 3xTF32 as a normal dtype and subsequent analysis compatible with this special dtype :-)