3 ms·
> It always made me wonder why liquid cooling wasn't more of a thing for datacenters. Liquid cooling is almost a defacto-standard in data centers in the HPC wo
by adev_ 3y ago
> It always made me wonder why liquid cooling wasn't more of a thing for datacenters.
Liquid cooling is almost a defacto-standard in data centers in the HPC world. The Top of the TOP500 machines are all liquid cooled. Not by choice, but due to physics constraints.
There is a big gap in power density between the HPC world and the usual datacenter-commodity-hardware world.
Commodity DS are designed with the assumption that the average machine will run with a fraction of it's maximum load. HPC systems at the opposite are designed to operate safely at 100% load all the time.
In a previous company where I worked, we attempted to install a medium size HPC cluster in a well-known commerical datacenter and network provider. The commercial of the DS almost felt from his chair when we announced the power requirements.
- bayindirh 3y ago> we attempted to install a medium size HPC cluster in a well-known commerical Datacenter and network provider. The commercial of the DS almost fall from his chair when we announced the power requirements. Heh. We tried it too. They didn’t believe that a single node used their entire rack’s budget at first.
- KolenCh 3y agoI'm going through a similar thing. Next week our procurement is arriving, but most likely majority of the nodes will sit in a box for a while before they can figure out how to power them... Also, this involves a UK university where it would takes forever to upgrade the power delivery to the building/floor/room. There's no planned upgrade whatsoever so people just need to make do with what we have.
- mrgaro 3y agoSounds fascinating. Can you give any more details? What kind of nodes are they and how they differ from "traditional" DC hardware, say from Supermicro?
- lhoff 3y agoThe difference is GPUs. A normal dual socket system serving a database or webserver use under medium load around 200-300W, One of these [1] equipped with 10xA100 can easily use in the ballpark of 3kW under load. So we are talking 10x the power usage. [1]https://www.supermicro.com/en/products/system/gpu/5u/sys-521ge-tnrt https://www.supermicro.com/en/products/system/gpu/5u/sys-521...
- mrgaro 3y agoThanks, this makes sense!
- deleted 3y ago[deleted]
- KolenCh 3y agoI’m a bit skeptical about the claim that the top of the TOP500 are all liquid cooled due to physical constraints. (Where the “physics” here seems to mean the density and property of air cool in general but ignoring the environment of the machine such as the ambient weather.) The one data point I know well is NERSC, and has been air cooled until relatively recently. Part of the success of air cool in the past was the great Bay Area weather that is typically cooled enough. I forgot the exact reason for upgrading to water cool, but it was before the recent upgrade to a new machine (Perlmutter). It may have to do with our weather getting more extreme so that there are more incidences that air cool only will “throttle” your machine. The reason I’m skeptical about that claim is that the ambient temperature also play a role. But certainly the density is increasing so may be the biggest supercomputers at the moment are too dense (in power consumption) to be air cooled.
- adev_ 3y agoIf you take the top5 of the top500 (01/2024): - Frontier: Liquid cooled - Aurora: liquid cooled - Eagle: liquid cooled - Fugaku: liquid cooled - Lumi: Liquid cooled with heat recycling. The need for liquid cooled arrived with the usage of accelerators that bumped significantly the power density. That said, Fugaku which is not accelerator driven is also liquid cooled today. I am surprised to learn that NERSC was air cooled until recently. Most of the BGQ machines, that were dominating the top500 10 years ago, were already liquid cooled.
- KolenCh 3y agoI probably wasn't clear. I don't doubt your first statement "The Top of the TOP500 machines are all liquid cooled." My skepticism is your second statement "due to physics constraints". It may very well be true but I'm just skeptical that it can't be done with a combination of engineering and ambient environment (mostly temperature but also a tradeoff between density and space). May be I focused on the word physics too much (being a physicist), but it seems the decision would basically be a cost-benefit analysis and risk management which involves many factors including money, maturity of solutions in the market, safety, etc. For example, at NERSC, there are real risk of a massive earthquake long overdue (the whole floor is quake-proof, but I guess the risk is reduced, not eliminated), so I guess they probably have considered this in that design choice. But perhaps you're right that the physics is the ultimate factor here: what comes in must goes out. With order of magnitude scale of increase in power coming in, water seems to be the most effective and cheap entity to absorb those heat and be transported outside the floor very efficiently.