6 ms·
What's the definition of "one supercomputer" for the purposes of TOP500? For example, why doesn't one of Google's warehouses qualify? Or the whole of Google, f
by improbable22 8y ago
What's the definition of "one supercomputer" for the purposes of TOP500?
For example, why doesn't one of Google's warehouses qualify? Or the whole of Google, for that matter. A bit of googling didn't find my anything very satisfactory.
- zeusk 8y agoor the hyperscale clouds (AWS, Azure, GCP).
- dooglius 8y agoI believe anything that can perform the LINPACK benchmark is eligible, though to qualify the owner of the computer would have to voluntarily run the benchmark and submit their results. Google has chosen not to submit any results, probably because they have better things to do with their warehouses than run benchmarks.
- stephencanon 8y agoBecause they don't submit results. You have to enter to win. Tangentially, Top500 results are based on one benchmark (latency of enormous double precision matrix triangular factorization), which is relatively far removed from what Google is optimizing for.
- jcranmer 8y agoTOP500 doesn't include distributed systems. Essentially, every computer on TOP500 is a single computer than you can log onto. By contrast, Google's data warehouse would qualify as a large cluster of individual systems. Note that not all supercomputers are on TOP500. Blue Waters is perhaps the most notable one to not bother reporting its performance (it would probably have been #1 had it done so when it came out, and today it would fall around 13th or so).
- dwyerm 8y agoI'm not sure that's true. At the very least, EC2 made a showing with C3 instances that made it to #64 in 2013. https://www.top500.org/system/178321 https://www.top500.org/system/178321
- bbgm 8y agoCC2 was #42, which was cool just because of the number. https://aws.amazon.com/blogs/aws/next-generation-cluster-computing-on-amazon-ec2-the-cc2-instance-type/ https://aws.amazon.com/blogs/aws/next-generation-cluster-com...
- throwaway2048 8y agoWhat you are talking about is called "single system image", that is a single address space, storage is globaly visible, etc. This concept is pretty much completely dead in the supercomputing world, and has been for decades.
- bbatha 8y agoOne thing google et all are missing from a typical super computer is infiniband style interconnects. They provide integrations with parallel data libraries like mpi and offer “3d” networking that will take into account physical distance between nodes and can do single rack mesh networking to avoid the overhead of switching. Despite google having lots of compute power they probably can’t leverage it in the way that the LINPACK benchmarks need.
- improbable22 8y agoThanks, this is interesting. It would be somehow satisfying if their LINPACK benchmarks would actually not be beaten by Google et al. (And their real workloads too.) But how tightly can you really connect 27000 GPUs? Would be curious if anyone has a more technical article handy about what's different.
- jcranmer 8y agoThe list of top supercomputers isn't a list of which systems have the most ALUs that you can shove floats through (though that is definitely a strong correlate). The difficult part in HPC is actually being able to keep those ALUs fed with floats. In large HPC applications, the communication is the principle bottleneck in being able to scale up [1]. Communication patterns for HPC application also tend to very much have a bursty everybody-is-sending-at-the-same-time pattern, which makes it very easy to saturate a typical star-like Ethernet network configuration (supercomputers typically use a torus or mesh-style interconnect). For GPUs, one trick you can do is to do GPU-to-GPU communication that bypasses the CPU. I don't believe the hardware that extends this to do CPU-less transfer systems across different nodes is common on non-HPC systems. [1] One of the main criticisms of LINPACK as a benchmark is that it is a low-communication benchmark. Essentially, you're doing O(n^3) computation on O(n^2) communication. In many benchmarks, such as grid simulation, the ratio of computation is communication is constant with respect to size.
- stephencanon 8y ago> One of the main criticisms of LINPACK as a benchmark is that it is a low-communication benchmark. Essentially, you're doing O(n^3) computation on O(n^2) communication. In many benchmarks, such as grid simulation, the ratio of computation is communication is constant with respect to size. This is critical and poorly communicated to most people outside the HPC world.
- ychen306 8y agoThe difference between a supercomputer and a data center is how "connected" the computations are; supercomputer optimizes the communication between nodes. To put it another way, a data center does a lot of work but, most of the time, for different applications (services) whose dependencies are "sparse".
- deepnotderp 8y agoIt's all in the Interconnects. The hard part of supercomputing is moving data, not computing.
- DannyBee 8y agoWhat makes you think google doesn't have good enough interconnects in their data centers? infiniband is not that impressive anymore "2009: of the top 500 supercomputers in the world, Gigabit Ethernet is the internal interconnect technology in 259 installations, compared with 181 using InfiniBand.[20] "
- dekhn 8y agoDefinition of supercomputer varies, but TOP500 is based on LINPACK. I am not aware of Google running any LINPACK benchmarks on their hardware (except maybe in Cloud VMs)? As others will say, the classic Google warehouses weren't really supercomputers, but more like massive clusters with a high cross-sectional bandwidth, but with very high latency, and they didn't run an MPI stack.