5 ms·
Interesting. What benchmark are you running? Getting it within 10% on such drastically different hardware looks wrong to me. And sounds like GPU is not a bottle
by dchichkov 7y ago
Interesting. What benchmark are you running? Getting it within 10% on such drastically different hardware looks wrong to me. And sounds like GPU is not a bottleneck in that particular case. Could it be that the system is CPU bound (with some horrid thread contention in TensorFlow or something)?
- bonoboTP 7y agoLambdaLabs' benchmark: https://github.com/lambdal/lambda-tensorflow-benchmark https://github.com/lambdal/lambda-tensorflow-benchmark It's certainly around 10-15% and it's not CPU bound, as the benchmark uses synthetic images without the need for data loading and preprocessing. I cannot rerun them now for precise numbers because my ROCm packages are broken and I don't have the spare hours to fix things. Surely it can be fixed but needs fiddling.
- dchichkov 7y agoYou can still be CPU bound, even if the benchmark is using synthetic images. A quick check could be to look if your 2080 Ti is 100% utilized (with nvidia-smi). And calculate percentage of theoretical peak FP32 throughput that you are getting for both cards. It is just that "more normal" result that you would expect to get, would be a drastic differences in performance. Note, Cvikli comment on "May 12" on that GitHub page: * The result with RNN networks on 1 Radeon VII and 1080ti was close to the same * Comparing convolutional performance the 4AMD and 4Nvidia, difference got really huge because of cuDNN for Nvidia cards. We can get more than 10x performance from the 1080Ti than the Radeon VII card. We find this difference in speed a little too big at image recognition cuDNN, I can't believe that this should happen and the hardware shouldn't be able to achieve the same. Note also in that comment, it looks like 4xGPU system just didn't work. It is hard to fully utilize GPU resources. Getting same result normally means that the same bottleneck is being hit. Algorithmic differences could result in drastic performance differences (10x). And as a final point of small things making a difference, note it is 1080Ti in that comment.
- bonoboTP 7y agoI fixed the installation and ran a benchmark again: https://news.ycombinator.com/item?id=21666411 https://news.ycombinator.com/item?id=21666411 The benchmark is not CPU bound, as evidenced by looking at rocm-smi, nvidia-smi (GPU usage at 100%) and htop (low CPU usage).
- dchichkov 7y agoNice. Your results could be more reproducible, if you'd include CuDNN version. Turing architecture support (and optimizations specific to DNN training) are still relatively recent and there are differences in performance between the versions [1]. Multiplying matrices is kinda tricky ;) [2] [1] https://developer.nvidia.com/cudnn https://developer.nvidia.com/cudnn . [2] https://scholar.google.com/scholar?as_ylo=2019&q=nvidia+gemm https://scholar.google.com/scholar?as_ylo=2019&q=nvidia+gemm
- bonoboTP 7y agoCannot edit that comment any more, but it's driver version 430, CUDA 10.1, CuDNN 7.5.
- dnautics 7y agoWe (lambda) are thinking about offering amd virtual workstation instances. One concern we had was that AMDs drivers have in the past had poor stability. Did you do burn in testing as well?
- bonoboTP 7y ago> Did you do burn in testing as well? Not sure what that means. I did this: https://news.ycombinator.com/item?id=21666411 https://news.ycombinator.com/item?id=21666411 I run the fan at 100% and even after several thousand iterations it keeps its speed at 274-275 im/sec and 68 C temperature with the stock fan and a standard desktop chassis kept open.
- dnautics 7y agoCool! Hey if you want to get together and chat about amd GPUs, I'm i@<lambda url>, happy to get you a coffee or beer or something.
- rrss 7y agoWhat's drastically different between Radeon VII and 2080 Ti? They have basically the same peak fp32 tflops.