4 ms·
Thank you for cross-referencing that, the data does look accurate and my statement now seems exaggerated. I do think we need skepticism on the A100 charts thoug
by tbenst 6y ago
Thank you for cross-referencing that, the data does look accurate and my statement now seems exaggerated. I do think we need skepticism on the A100 charts though until third party benchmarks.
> In my experience NVidia benchmark numbers in deep learning are rarely lies - they are highly optimised, in optimal conditions and rarely achievable in the real world.
Right, but Nvidia claimed 1360 images/sec for resnet-50 on imagenet. To my knowledge this still hasn’t been realized by a third party. It also isn’t a 4x improvement for fp32 vs fp16, that’s comparing to previous generation. Improvement is more like 1.5x: https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v100-vs-titan-v-vs-1080-ti-benchmark/ https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v...
Even in very simple synthetic benchmarks the speed up is only 2x: https://github.com/tensorflow/benchmarks/issues/77#issuecomment-349838985 https://github.com/tensorflow/benchmarks/issues/77#issuecomm...
I have not seen any benchmarks showing an 8x speedup. Have you? If not -> Nvidia lied.
- nl 6y ago> Right, but Nvidia claimed 1360 images/sec for resnet-50 on imagenet. On https://images.nvidia.com/content/technologies/volta/pdf/volta-v100-datasheet-update-us-1165301-r5.pdf https://images.nvidia.com/content/technologies/volta/pdf/vol... they claim 1,525 images/sec (!) Dell hit 5,243 images/sec with one of their 4 V100s servers, which comes to 1,310 images/sec per V100. I find it very believable that NVidia would get ~200 images/sec more, since Dell jumped 50% with a change in their CPU/GPU connection topology. See https://www.dell.com/support/article/en-au/sln317397/deep-learning-performance-on-v100-gpus-with-resnet-50-model?lang=en https://www.dell.com/support/article/en-au/sln317397/deep-le...
- nl 6y agoAlso I just noticed that Google's XLA gets 1278 images/second on a single V100 in FP16 mode. https://www.tensorflow.org/xla https://www.tensorflow.org/xla
- tbenst 6y agoThank you for finding that! I’m glad to see the situation has improved since I last looked. 3x improvement over fp32 is impressive for sure. Their marketing claims of 8x still bother me though.
- nl 6y agoWhat exactly is the 8x claim? The thing I've seen is very limited (NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers.[1]) which is probably true. The problem with performance improvements is the diminishing returns part of Amdahl Law: an 8x improvement in math performance will just mean the math part becomes less important in terms of absolute performance. In any case, I've found NVidia's claims in the machine learning area to be pretty good. Like most claims you have to read carefully to see exactly what the claim is, but that's not uncommon with performance claims. [1] https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html https://docs.nvidia.com/deeplearning/performance/mixed-preci... [2] https://en.wikipedia.org/wiki/Amdahl%27s_law#Relation_to_the_law_of_diminishing_returns https://en.wikipedia.org/wiki/Amdahl%27s_law#Relation_to_the...
- tbenst 6y ago> NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers Right but I’ve benchmarked the best case scenario, ie a large GEMM call in C++, and still not seen anywhere close to 8x. I’ve never seen a code example, no matter how limited, showing a 8x speed up.
- llukas 6y agohttps://developer.nvidia.com/blog/programming-tensor-cores-cuda-9/ https://developer.nvidia.com/blog/programming-tensor-cores-c... See cuBLAS mixed-precision GEMM.
- deleted 6y ago[deleted]