3 ms·
> NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers Right but I’ve b
by tbenst 6y ago
> NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers
Right but I’ve benchmarked the best case scenario, ie a large GEMM call in C++, and still not seen anywhere close to 8x. I’ve never seen a code example, no matter how limited, showing a 8x speed up.
- llukas 6y agohttps://developer.nvidia.com/blog/programming-tensor-cores-cuda-9/ https://developer.nvidia.com/blog/programming-tensor-cores-c... See cuBLAS mixed-precision GEMM.