3 ms·
Thanks, this is interesting. It would be somehow satisfying if their LINPACK benchmarks would actually not be beaten by Google et al. (And their real workloads
by improbable22 8y ago
Thanks, this is interesting. It would be somehow satisfying if their LINPACK benchmarks would actually not be beaten by Google et al. (And their real workloads too.)
But how tightly can you really connect 27000 GPUs? Would be curious if anyone has a more technical article handy about what's different.
- jcranmer 8y agoThe list of top supercomputers isn't a list of which systems have the most ALUs that you can shove floats through (though that is definitely a strong correlate). The difficult part in HPC is actually being able to keep those ALUs fed with floats. In large HPC applications, the communication is the principle bottleneck in being able to scale up [1]. Communication patterns for HPC application also tend to very much have a bursty everybody-is-sending-at-the-same-time pattern, which makes it very easy to saturate a typical star-like Ethernet network configuration (supercomputers typically use a torus or mesh-style interconnect). For GPUs, one trick you can do is to do GPU-to-GPU communication that bypasses the CPU. I don't believe the hardware that extends this to do CPU-less transfer systems across different nodes is common on non-HPC systems. [1] One of the main criticisms of LINPACK as a benchmark is that it is a low-communication benchmark. Essentially, you're doing O(n^3) computation on O(n^2) communication. In many benchmarks, such as grid simulation, the ratio of computation is communication is constant with respect to size.
- stephencanon 8y ago> One of the main criticisms of LINPACK as a benchmark is that it is a low-communication benchmark. Essentially, you're doing O(n^3) computation on O(n^2) communication. In many benchmarks, such as grid simulation, the ratio of computation is communication is constant with respect to size. This is critical and poorly communicated to most people outside the HPC world.
- bbatha 8y ago> But how tightly can you really connect 27000 GPUs? Not all that well currently, NVidia and others are working on GPU specific interconnects[0] but they don't have anywhere near the scale of traditional interconnects which have supported hundreds of thousands of nodes by the late 90s. On of the big challenges in modern super computer programming is in fact keeping the GPUs hot, which can often mean offloading work that needs high memory usage to CPUs. Unfortunately my knowledge here is a little dated, I interned at Los Alamos National Lab from 2008 - 2012 when they were doing a lot of rearchitecting of old codes for RoadRunner, the first peta-scale computer. It used Cell chips in accelerator cards and predicated a lot of the challenges in GPU programming, but did not fully elucidate them. For instance we didn't have CUDA! If I had to take my guess the first exa-scale computer is going to be the one that solves the GPU interconnect problem at scale. 0: https://www.nvidia.com/en-us/data-center/nvlink/ https://www.nvidia.com/en-us/data-center/nvlink/
- shaklee3 8y agoExactly. First there was nvswitch, which dramatically increased the bandwidth over pcie. But that didn't scale to a large number of GPUs. Then there was nvswitch, which solved the scaling problem inside a node. I wouldn't be surprised if the next leap is something like nvlink cables between nodes that don't need traditional routing capabilities.
- shaklee3 8y agoFirst there was nvlink, rather.
- greglindahl 8y agoRoadrunner was... special... in that regard, requiring even more effort and bizarreness than the typical HPC GPU setup. I remember that the pre-install plan for Roadrunner Linpack was a 50 page document. Also, it's worth noting that GPU HPC computing was already in full swing around the same time: CUDA was first released in June, 2007, which is the same time that Roadrunner released its first Top500 entry.