4 ms·
I appreciate the time and care that went into this post, and there’s a nice discussion of various features. Unfortunately the performance charts are completely
by tbenst 6y ago
I appreciate the time and care that went into this post, and there’s a nice discussion of various features.
Unfortunately the performance charts are completely devoid from reality, and in particular the discussion on tensorcores may be true from an instruction count perspective but does not reflect any third-party benchmark I’ve seen. For example: https://lambdalabs.com/blog/2080-ti-deep-learning-benchmarks/ https://lambdalabs.com/blog/2080-ti-deep-learning-benchmarks.... Nvidia has a history of straight-up-lying about tensorcore and other benchmarks (for example, see this thread from right after Nvidia announced an 8x improvement in speed on imagenet in tensorflow for V100: https://github.com/tensorflow/benchmarks/issues/77 https://github.com/tensorflow/benchmarks/issues/77)
In general, fp16 is only 30-40% faster than fp32, and occasionally 2x in really optimal conditions.
- llukas 6y agoCannot really lie here: https://mlperf.org/ https://mlperf.org/
- YetAnotherNick 6y agoI don't see the timing comparison between FP16 and FP32. Can you point me at it if it is there.
- nl 6y ago> Unfortunately the performance charts are completely devoid from reality, and in particular the discussion on tensorcores may be true from an instruction count perspective but does not reflect any third-party benchmark I’ve seen. For example: https://lambdalabs.com/blog/2080-ti-deep-learning-benchmarks/ https://lambdalabs.com/blog/2080-ti-deep-learning-benchmarks... The performance numbers posted here appear to almost exactly reflect the LambdaLabs numbers. Lambda Labs: the RTX 2080 Ti is 96% as fast as Titan V, 73% as fast as Tesla V100 (32 GB) timdettmers: RTX 2080 Ti normalized to 1, Titan V looks about 1.1 to 1.2, V100 is just below 1.5 > for example, see this thread from right after Nvidia announced an 8x improvement in speed on imagenet in tensorflow for V100 Well a non-NVidia person managed to get it up to just above 4x improvement without "using unreleased libraries from NVIDIA". From the same thread: https://github.com/tensorflow/benchmarks/issues/77#issuecomment-394541623 https://github.com/tensorflow/benchmarks/issues/77#issuecomm... In my experience NVidia benchmark numbers in deep learning are rarely lies - they are highly optimised, in optimal conditions and rarely achievable in the real world. About what you'd expect from a vendor benchmark.
- tbenst 6y agoThank you for cross-referencing that, the data does look accurate and my statement now seems exaggerated. I do think we need skepticism on the A100 charts though until third party benchmarks. > In my experience NVidia benchmark numbers in deep learning are rarely lies - they are highly optimised, in optimal conditions and rarely achievable in the real world. Right, but Nvidia claimed 1360 images/sec for resnet-50 on imagenet. To my knowledge this still hasn’t been realized by a third party. It also isn’t a 4x improvement for fp32 vs fp16, that’s comparing to previous generation. Improvement is more like 1.5x: https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v100-vs-titan-v-vs-1080-ti-benchmark/ https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v... Even in very simple synthetic benchmarks the speed up is only 2x: https://github.com/tensorflow/benchmarks/issues/77#issuecomment-349838985 https://github.com/tensorflow/benchmarks/issues/77#issuecomm... I have not seen any benchmarks showing an 8x speedup. Have you? If not -> Nvidia lied.
- nl 6y ago> Right, but Nvidia claimed 1360 images/sec for resnet-50 on imagenet. On https://images.nvidia.com/content/technologies/volta/pdf/volta-v100-datasheet-update-us-1165301-r5.pdf https://images.nvidia.com/content/technologies/volta/pdf/vol... they claim 1,525 images/sec (!) Dell hit 5,243 images/sec with one of their 4 V100s servers, which comes to 1,310 images/sec per V100. I find it very believable that NVidia would get ~200 images/sec more, since Dell jumped 50% with a change in their CPU/GPU connection topology. See https://www.dell.com/support/article/en-au/sln317397/deep-learning-performance-on-v100-gpus-with-resnet-50-model?lang=en https://www.dell.com/support/article/en-au/sln317397/deep-le...
- sytelus 6y ago> fp16 is only 30-40% faster than fp32 It really depends on workload. For ImageNet and resnet type architecture its not unusual to get 3X speed up. It also depends on if you do full fp16, leave BNs alone etc. This is a lot because instead of training for 6 days, you now training for 2 days.
- solidasparagus 6y ago> ImageNet and resnet type architecture its not unusual to get 3X speed up Source on this? I've done a good bit of CV benchmarking work and I don't recall anything like a 3x boost. 30-40% improvement is much more in line with what I remember.
- timdettmers 6y agoI should have been a bit clearer what went into the charts. I do not use theoretical marketing numbers but real-life benchmark data from NVIDIA and 4 other sources of benchmark between Titan V, V100, RTX 2080, RTX 2080 Ti and Titan RTX. Since I calibrate a model that needs to satisfy all sources as best as it can I think the numbers are pretty accurate.
- tbenst 6y agoThank you for clarifying! I’m still skeptical of the chart’s A100 values but appreciated your reasonable attempt to de-bias. It’s always easier to critique then create so I also want to make sure I complement you on an excellent article :).
- timdettmers 6y agoThank you, I just updated the blog post with more detailed clarification of where the data comes from. One thing that I am quite sure of for the A100 is its transformer performance. It turns out, large transformers are so strongly bottlenecked by memory bandwidth that you can just use memory bandwidth alone to measure performance — even across GPU architectures. The error between Volta and Turning with a pure bandwidth model is less than 5%. The NVIDIA transformer A100 benchmark data shows similar scaling. So I am pretty confident on the transformer numbers. The computer vision numbers are more dependent on the network and it is difficult to generalize across all CNNs. For example, group convolution or depth-wise separable convolution based CNNs do not scale well with better GPUs and speedups will be small (1.2 - 1.5x) whereas some other networks like ResNet get pretty straightforward improvements (1.6x-1.7x). So CNN values are less straightforward because there is more diversity between CNNs compared to transformers.