31 ms·
5 petaflops in DGX? That alone would put one of those babies into TOP500 top 50, and a superPOD would make no. 1, no? Well, if it could do that performance on L
by Keyframe 6y ago
5 petaflops in DGX? That alone would put one of those babies into TOP500 top 50, and a superPOD would make no. 1, no? Well, if it could do that performance on Linpack/Rmax.
- jl2718 6y agoTop500 is generally based on double precision (64 bit).
- Miraste 6y agoLess than that in double precision. If my math is right, they manage around 0.4 petaflops. Their previous-gen SuperPOD was #22. If they do indeed add four of the new ones that will still put it at #1, with something like 230 petaflops.
- jabl 6y agoA previous generation superPOD (with V100 GPU's) is currently at #20. Watch this space...
- namibj 6y agoYeah, use this with PCIe4 and add a dense PCIe facric. Broadcom https://www.broadcom.com/products/pcie-switches-bridges/expressfabric/gen4/pex88096 https://www.broadcom.com/products/pcie-switches-bridges/expr... and Microchip https://www.microchip.com/wwwproducts/en/PM42100 https://www.microchip.com/wwwproducts/en/PM42100 have 98/100 lane PCIe4 fabric switches to offer.
- sabalaba 6y agoNo, the A100 has a 19.5 TFLOP theoretical peak for SGEMM[1], real world benchmarks will likely achieve 93% of that, and so the DGX A100 will be 145 TFlops of FP32 SGEMM performance or 0.145 FP32 PFLOPS. Maybe in 72 FP64 TFLOPS. FP64 is what the TOP500 benchmaks count.[2] The 5 "petaflops" number is a creatively constructed marketing number based on FP16 TensorCore "flops", sparse matrix calculations, and then multiplying by 8x for some reason. They basically take the 19.5 FP32 TFLOPS number and multiply it by 32x to get to the claimed 624 "TFLOPS" for a single A100. 8 * 624 = 5 "petaflops". I see they get 2x by actually using FP16 instead of FP32, 2x from counting sparse matrix ops as dense ops, and 8x from somewhere else that I have no idea. [1] https://devblogs.nvidia.com/nvidia-ampere-architecture-in-depth/ https://devblogs.nvidia.com/nvidia-ampere-architecture-in-de... [2] https://www.top500.org/resources/frequently-asked-questions/ https://www.top500.org/resources/frequently-asked-questions/
- deleted 6y ago[deleted]
- rrss 6y ago> 8x from somewhere else that I have no idea. 8 GPUs in the box.
- sabalaba 6y agoNo, I already multiplied 624 TFLOPS / GPU * 8 GPU = 4992 TFLOPS (the 5 petaflops number). I'm saying that you are still missing another 8x on the way from 19.5 TFLOPS / GPU to 624 TFLOPS / GPU. 19.5 (base FP32 theoretical peak performance) * 2 (FP16 instead of FP32) * 2 (counting sparse matrix ops as dense ops) * 8 (unknown) = 624 TFLOPS.
- rrss 6y agoFP16 tensorcore = 312 tflops x 2 (counting sparse as dense) = 624 tflops x 8 GPUs = 5 "pflops" The missing 8x you are looking for is just because tensorcore math is much faster than their normal fma path.