4 ms·
Thank you for finding that! I’m glad to see the situation has improved since I last looked. 3x improvement over fp32 is impressive for sure. Their marketing cla
by tbenst 6y ago
Thank you for finding that! I’m glad to see the situation has improved since I last looked. 3x improvement over fp32 is impressive for sure. Their marketing claims of 8x still bother me though.
- nl 6y agoWhat exactly is the 8x claim? The thing I've seen is very limited (NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers.[1]) which is probably true. The problem with performance improvements is the diminishing returns part of Amdahl Law: an 8x improvement in math performance will just mean the math part becomes less important in terms of absolute performance. In any case, I've found NVidia's claims in the machine learning area to be pretty good. Like most claims you have to read carefully to see exactly what the claim is, but that's not uncommon with performance claims. [1] https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html https://docs.nvidia.com/deeplearning/performance/mixed-preci... [2] https://en.wikipedia.org/wiki/Amdahl%27s_law#Relation_to_the_law_of_diminishing_returns https://en.wikipedia.org/wiki/Amdahl%27s_law#Relation_to_the...
- tbenst 6y ago> NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers Right but I’ve benchmarked the best case scenario, ie a large GEMM call in C++, and still not seen anywhere close to 8x. I’ve never seen a code example, no matter how limited, showing a 8x speed up.
- llukas 6y agohttps://developer.nvidia.com/blog/programming-tensor-cores-cuda-9/ https://developer.nvidia.com/blog/programming-tensor-cores-c... See cuBLAS mixed-precision GEMM.