3 ms·
Would be curious to see your benchmarks. Btw, Nvidia will be providing support for fp8 in a future release of CUDA - https://github.com/NVIDIA/TransformerEngine
by varunkmohan 4y ago
Would be curious to see your benchmarks. Btw, Nvidia will be providing support for fp8 in a future release of CUDA - https://github.com/NVIDIA/TransformerEngine/issues/15 https://github.com/NVIDIA/TransformerEngine/issues/15
I think TMA may not matter as much for consumer cards given the disproportionate amount of fp32 / int32 compute that they have.
Would be interesting to see how close to theoretical folks are able to get once CUDA support comes through.
- touisteur 4y agoWell every since I've read these papers doing FFTs faster than cuFFT using tensor cores (although FFT isn't supposed to be helped by more flops but only better memory bandwidth) and also fp32-level accuracy on convolutions with 3x tf32 tensor cores sweeps (available in cutlass) I'm quite ready to believe some hype about TMA. Anything improving memory bandwidth for any compute part of the GPU is welcome. Also I'd like for someone to crack open RT cores and get the ray-triangle intersection acceleration out of Optics. Have you seen the FLOPS on these things?
- touisteur 4y agoBefore I forget, here's the link to the tcFFT paper https://ar5iv.labs.arxiv.org/html/2104.11471v1 https://ar5iv.labs.arxiv.org/html/2104.11471v1 And the fp32-gemm-with-tf32 https://arxiv.org/abs/2203.03341 https://arxiv.org/abs/2203.03341