3 ms·
If I understand your comment correctly, we're taking a stable but not that relevant metric, because the real players of the market are too secretive, fast and f
by makeitdouble 2y ago
If I understand your comment correctly, we're taking a stable but not that relevant metric, because the real players of the market are too secretive, fast and far ahead to allow for simple comparisons.
From a distance, it kinda sounds like listening to kids brag about their allowance while the adults don't want to talk about their salary, and try to draw wider conclusions from there.
- wbl 2y agoEven the DoE posts top 500 results when they commission a supercomputer.
- makeitdouble 2y agoDoE has absolutely no incentive (nor need, I'd argue) to compare their supercomputers to commercially owned data center operations though. Comparing their crazy expensive custom built HPC to massive arrays of customer grade hardware doesn't bring them additional funds, nor help them more PR wise than being the owner of the fastest individual clusters. Being at the top of some heap is visibly one of their goal: https://www.energy.gov/science/high-performance-computing https://www.energy.gov/science/high-performance-computing
- khm 2y agoDOE clusters are also massive arrays of customer grade hardware. Private cloud can only keep up in low precision work, and that is why they're still playing with remote memory access over TCP, because it's good enough for web and ML. High precision HPC exists in the private cloud, but you only hear "we don't want to embarrass others" excuses because otherwise you would be able to calculate the cost. On prem HPC is still very, very much cheaper than hiring out.
- zekrioca 2y agoIt seems there was a misunderstanding, as I haven't made any value judgment about LINPACK. Yes, LINPACK is indeed "old" with a heavy focus on compute power. However, its simplicity serves as a reliable baseline for the types of workflows that supercomputers are designed to handle. Also, at their core, most AI workloads perform essentially the same operations as HPC, albeit with less stability—which, I admit, is a feature, but likely the reason AI-focused systems do not prioritize LINPACK as much. I am simply saying that any useful metric needs to not only be "stable", but also simple to grasp. Take Green500, probably a significant benchmark for understanding how algorithms consume power, but "too complex" to explain: yet, many cloud providers with their AI supercomputers avoid competing against HPC supercomputers in this domain. This avoidance isn’t necessarily due to secrecy but rather inefficiencies inherent to cloud systems. Consider PUE (Power Usage Effectiveness)—a highly misleading metric that cloud providers frequently tout. PUE can easily be manipulated, especially with the use of liquid cooling, which is why optimizing for it has become a major factor contributing to water disruptions in several large cities worldwide.