6 ms·
It seems odd to focus on the CPU here, when presumably the vast majority of those flops re coming from the NVIDIA parts? Are the CPUs expected to contribute si
by owlbite 3y ago
It seems odd to focus on the CPU here, when presumably the vast majority of those flops re coming from the NVIDIA parts?
Are the CPUs expected to contribute significant compute, as opposed to marshaling data in/out of the real compute units?
- jefft255 3y agoYes, CPUs are still the main workhorse for many scientific workloads. Sometimes just because the code hasn’t been ported, sometimes because it’s just not something that a GPU can do well.
- londons_explore 3y ago> just because the code hasn’t been ported, Seems stupid to use millions of dollars of supercomputer time just because you can't be bothered to get a few phd students to spend a few months rewriting in CUDA...
- cmdrk 3y agosometimes the code is deeply complex stuff that has accumulated for over 30 years. to _just_ rewrite it in CUDA can be a massive undertaking that could easily produce subtly incorrect results that end up in papers could propagate far into the future by way of citations etc
- throwaway10965 3y agoSounds like a great job for LLMs. Are there any public repositories of this code? I want to try.
- mlyle 3y agoSounds like a -terrible- job for LLMs, because this is all about attention to detail. Order of operations and specific constructs of how floating point work in the codes in question are usually critical. Have fun: https://www.qsl.net/m5aiq/nec-code/nec2-1.2.1.2.f https://www.qsl.net/m5aiq/nec-code/nec2-1.2.1.2.f
- throwaway10965 3y agoAttention to detail can come later when there's something that humans can get started with. I did not mean that LLM could do it all alone.
- mlyle 3y agoA human has to have the knowledge of what the code is trying to do and what the requisites are for accuracy and numerical stability. There's no substitute for that. Having a translation aid doesn't help at all unless it's perfect: it's more work to verify the output from a flawed tool than to do it right in this case.
- londons_explore 3y agoAll the more reason to rewrite it... You don't want some mistake in 30 year old COBOL code to be making your 2023 experiment to have wrong results.
- gmueckl 3y agoThat's the complete opposite of what is actually the case: some of that really old code in these programs is battle-tested and verified. Any rewrite of such parts would just destroy that work for no good reason.
- mlyle 3y agoThe whole point is in these older numerical codes is that they're proven and there's a long history of results to compare against.
- dpe82 3y ago*FORTRAN.
- _a_a_a_ 3y agoWhy don't YOU take some old code and rewrite it. I tried it for some 30+ year old HPC code and it was a grim experience and I failed hard. So why not keep your lazy, fatuous suggestions to yourself.
- mlyle 3y agoA supercomputer might cost $200M and use $6M of electricity per year. Amortizing the supercomputer over 5 years, a 12 hour job on that supercomputer may cost $63k. If you want it cheaper, your choices are: A) run on the supercomputer as-is, and get your answer in 12 hours (+ scheduling time based on priority) B) run on a cheaper computer for longer-- an already-amortized supercomputer, or non-supercomputing resources (pay calendar time to save cost) C) try to optimize the code (pay human time and calendar time to save cost) -- how much you benefit depends upon labor cost, performance uplift, and how much calendar time matters. Not all kinds of problems get much uplift from CUDA, anyways.
- jonwachob91 3y ago>> A supercomputer might cost $200M and use $6M of electricity per year. I'm curious, what university has a $200MM super computer? I know governments have numerous Supercomputers that blow past $200MM in build price, but what universities do?
- sophacles 3y agoUniversity of Illinois had Blue Waters ($200+MM, built in ~2012, decomissioned in the last couple years). https://www.ncsa.illinois.edu/research/project-highlights/blue-waters/ https://www.ncsa.illinois.edu/research/project-highlights/bl... https://en.wikipedia.org/wiki/Blue_Waters https://en.wikipedia.org/wiki/Blue_Waters They have always had a lot of big compute around.
- mlyle 3y ago> I know governments have numerous Supercomputers that blow past $200MM in build price, but what universities do? Even when individual universities don't-- governments have supercomputing centers that universities are a primary user of and often charge back value of computing time to the university or it is a separate item that is competitively granted. Here we're talking about Jupiter, which is a ~$300M supercomputer where research universities will be a primary user.
- otabdeveloper4 3y agoCUDA is buggy proprietary shit that doesn't work half the time or segfaults with compiler errors. Basically, unless you have a very specific workload that NVidia has specifically tested, I wouldn't bother with it.
- brnt 3y agoThe JSC employs a good number of people doing exactly this.
- bee_rider 3y ago>> just because the code hasn’t been ported, sometimes because it’s just not something that a GPU can do well. > Seems stupid to use millions of dollars of supercomputer time just because you can't be bothered to get a few phd students to spend a few months rewriting in CUDA... Rewriting code in CUDA won’t magically make workloads well suited to GPGPU.
- wang_li 3y agoIt's highly likely that a workload that is suitable to run on hundreds of disparate computers with thousands of CPU cores is going to be equally well suited for running on tens of thousands of GPU compute threads.
- atq2119 3y agoNot necessarily. GPUs simply aren't optimized around branch-heavy or pointer-chasey code. If that describes the inner loop of your workload, it just doesn't matter how well you can parallelize it at a higher level, CPU cores are going to be better than GPU cores at it.
- monocasa 3y agoThey're not that disparate; the workloads are normally very dependent on the low latency interconnect of most supercomputers.
- hulitu 3y agoCUDA ? I thought rust was the future. /s
- xadhominemx 3y agoAgreed, the CPUs are not performing the scientific calculations in this system. Also note — this project is quite modest in scale. Dozens of GenAI clusters larger than this computer will be installed at cloud data centers in the next 18 months.
- hexane360 3y agoI don't know how you can describe "equal to the world's fastest supercomputer, which was built less than a year ago" as "quite modest".
- RetroTechie 3y ago"The Jupiter will instead have SiPearl’s ARM processor based on ARM’s Neoverse V1 CPU design. SiPearl has designed the Rhea chip to be universally compliant with many accelerators, and it supports high-bandwidth memory and DDR5 memory channels." And "Jülich is also building out its machine-learning and quantum computing infrastructure, which the supercomputing center hopes to plug in as accelerator modules hosted at its facility." So a modular setup, where different aspects can be upgraded as needed. Btw: > Also note — this project is quite modest in scale. "Exascale" and €273M doesn't sound modest to me. No matter what it's compared against.
- pclmulqdq 3y agoAlmost all of the AI computers being built now are relatively modestly sized compared to a supercomputer. All but the biggest ones are at or under the low hundreds of nodes (low thousands of GPUs). The only real exceptions are the few AI hyperscale companies that want to sell GPU computing to others.
- bee_rider 3y agoDo AI hyperscalers devote a their whole system to one big run anyway? If they don’t, then those are big clusters in the sense that AWS is the world’s biggest supercomputer, which is to say, not.
- brucethemoose2 3y agoSIMD heavy CPUs can provide quite respectable HPC throughout. The US Dept of Energy had very favorable things to say about the Fujitsu A64FX, which is architecturally similar to the SiPearl Rhea (HBM memory, ARM SVE happy, fast interconnect): https://www.osti.gov/biblio/1965278 https://www.osti.gov/biblio/1965278 They seemed to like the easy porting and flexible programming (since its "just" CPU SIMD) and specifically describe it as competitive with Nvidia: > To highlight, the pink line represents the energy efficiency metric for A64FX in boost power mode (described in Section IV-C) with an estimated TDP of 140 W and surpassed by the red and yellow lines that represent data for the Volta V100 GPU (highest) and KNL, respectively. The A64FX architecture scores better with the energy efficiency metric relative to the performance efficiency metric due to its low power consumption. In fact, ARM A64FX supercomputers topped the Green500 for some time, which is the global supercomputer power efficiency ranking, outclassing Nvidia/Intel/AMD machines.
- monocasa 3y agoYou'd be surprised. A lot of supercomputers aren't that much about individual CPU core perf, but having a lot of low power cores connected in a novel way. The BlueGene supercomputers were composed of low spec PowerPC cores (even for the time). High perf/watt matter more than just high perf/node, but even that balanced against 'how low latency can the interconnect be'. You then hit the high FLOP count with tons of nodes. To be fair Nvidia realized this paradigm years ago too, which is why they bought Mellanox.
- fsckboy 3y agoI can believe what you are saying, but what are you saying, is it heat (ie cooling), power (cost), or hw cost/flop (because these chips are cheaper than screamers) that makes this an optimal solution?
- monocasa 3y agoI'd describe the end goal as TCO of achievable FLOP. So a lot of factors come out of that, and a lot of designs that take interesting stabs at new balances towards that goal.
- cmdrk 3y agoit really depends on the workload. as other posters have said, not everything can/should be ported to GPU. some scientific calculations are simply not parallelizable in that way. typically at least in the US there's a mix of GPU-focused machines as well as traditional CPU-focused machines. the leadership class machines (i.e., the machines funded to push the FLOPS records) tend to be highly focused on GPU. one reason is fixed cooling/power availability. I assume these facilities are looking at ARM as a way to save 10-20% on power and thus cram that much more into the facility.