26 ms·
AMD-powered Frontier supercomputer breaks the exascale barrier
- uniqueuid 4y agoSince they are using AMD's accelerators as well [1], I do wonder whether any usage of these will trickle down and give us improvements in ROCm. Surely the people at these labs will want to run ordinary DL frameworks at some point - or do they have the money and time to always build entirely custom stacks? [1] AMD Instinct MI250x in this case.
- dragontamer 4y ago> Surely the people at these labs will want to run ordinary DL frameworks at some point I don't know about that. A lot of these labs are doing physics simulations and are probably happy to stick with their dense-matrix multiply / BLAS routines. Deep learning is a newer thing. These national labs can run them of course, but these national labs have existed for many decades and have plenty of work to do without deep learning. > or do they have the money and time to always build entirely custom stacks? Given all the talk about OpenMP compatibility and Fortran... my guess is that they're largely running legacy code in Fortran. Perhaps some new researchers will come in and try to get some deep-learning cycles in the lab and try something new.
- deleted 4y ago[deleted]
- marcosdumay 4y ago> Given all the talk about OpenMP compatibility and Fortran... my guess is that they're largely running legacy code in Fortran. The must used linear algebra library is written in Fortran. There's nothing "legacy" about it, it's just that nobody was able to replicate its speed in C.
- jcranmer 4y ago> The must used linear algebra library is written in Fortran. My understanding is that most supercomputers have the vendor provide their implementation of BLAS (e.g., if it's Intel-based, you're getting MKL) that's specifically tuned for that hardware. And these implementations stand a decent chance of being written in assembly, not Fortran.
- bee_rider 4y agoUsually C or Fortran superstructure, and assembly kernels. The clearest form of this is in BLIS, which is a C framework you can drop your assembly kernel into, and then it makes a BLAS (along with some other stuff) for you. But the idea is also present in OpenBlas. Lots of this is due to the legacy of gotoBlas (which was forked into OpenBlas, and partially inspired BLIS), written by the somewhat famous (in HPC circles at least) Kazushige Goto. He works at Intel now, so probably they are doing something similar.
- dragontamer 4y agoBLAS itself has been rewritten in Nvidia CUDA and AMD HIP, and is likely the workhorse in this case. (Remember that Frontier is mostly GPUs and the bulk of code should be GPU compatible) Presumably that old Fortran code has survived many generations of ports: Connection Machine, DEC Alpha, Intel Itanium, SPARC and finally today's GPU heavy systems. The BLAS layer keeps getting rewritten but otherwise the bulk of the simulators still works.
- nspattak 4y agoIf you are talking about netlib blas/lapack I am very confused by what you are saying because the fastest blas/lapack implementations are in c/c++.
- paulmd 4y agoI don't remember the exact specifics, but Fortran disallows some of the constructs that C/C++ struggle with aliasing on, so Fortran can often be (safely) optimized to much higher-performance code because of this limitation/knowledge. Like, it's always seemed like there's a certain amount of fatalism around Undefined Behavior in C/C++, like this is somehow how it has to be to write fast code but... it's not. You can just declare things as actually forbidden rather than just letting the compiler identify a boo-boo and silently do whatever the hell it wants. Of course it's not the right tool for every task, I don't think you'd write bit-twiddling microcontroller stuff in fortran, or systems programming. But for the HPC space, and other "scientific" code? Fortran is a good match and very popular despite having an ancient legacy even by C/C++ standards (both have, of course, been updated through time). Little less flexible/general, but that allows less-skilled programmers (scientists are not good programmers) to write fast code without arcane knowledge of the gotchas of C/C++ compiler magic.
- jcranmer 4y agoFrom my limited exposure to the HPC groups at the labs, there's a mixture of languages in use. It seems that modern C++ is the dominant language for a lot of new projects--some of the people I talked to were working on libraries that aggressively used C++11/C++14 features. The biggest challenge the national labs face is that there's not really any budget (or appetite) to rewrite software to take advantage of hardware features (particularly the GPU-based accelerator that's all the rage nowadays). You might be able to get a code rewritten once, but an era where every major HPC hardware vendor wants you to rewrite your code into their custom language for their custom hardware results in code that will not take advantage of the power of that custom hardware. OpenMP, being already fairly widespread, ends up becoming the easiest avenue to take advantage of that hardware with minimal rewriting of code (tuning a pragma doesn't really count as rewriting).
- Symmetry 4y agoAlso, while NVidia has been adding extra AI acceleration to their chips AMD has been throwing in extra double precision resources that HPC generally requires. If you're training an AI rather than simulating the climate/a thermonuclear explosion/etc then you're probably better off using NVidia cards but AMD made the right technical investments to get these supercomputer contracts.
- dekhn 4y agoIt's kind of surprising that nvidia hasn't purchased AMD. It really feels like there's a single company between the two that would be truly effective- AMD for the classic CPU oomph, nvidia for the GPU oomph, combining their strengths in interconnects. It would be a player from the high-end PC to the supercomputer market, without even pretending to go for the low-power market (ARM).
- jcranmer 4y ago> It's kind of surprising that nvidia hasn't purchased AMD. One word: antitrust. The discrete GPU market these days consists of Nvidia and AMD, with Intel only just now dipping its toes into the market (I don't think there's anything saleable to retail customers yet). Nvidia buying AMD would make it a true monopoly in that market, and there's no way that would pass antitrust regulators. Nvidia recently tried to buy ARM, and even that transaction was enough for antitrust regulators to say no.
- paulmd 4y agoI think based on recent history you can argue that NVIDIA is very aware of the potential anticompetitive actions that could result if they kill or even substantially pass AMD. There really used to be a lot of intra-generational tweaking and refinement, like if you look back at Maxwell there were really at least 3 and I suspect 4 total steppings of the maxwell architecture (GM107, GM204/GM200, and GM206 - and I suspect GM200 was a separate "stepping" too due to how much higher it clocks than GM204 - which is the opposite of what you'd expect from a big chip). Kepler had at least 4 major versions (GK1xx, GK110B, GK2xx, GK210), Fermi had at least 2 (although that's where I'm no longer super familiar with the exact details). Anyway point is there used to be a lot more intra-generational refinement, and I think that has largely stopped, it's just thrown over the wall and done. And I think the reason for that is that if NVIDIA really cranked full-steam ahead they'd be getting far enough ahead of AMD to potentially start raising antitrust concerns. We are now in the era of "metered performance release", just enough to stay ahead of AMD but not enough to actually raise problems and get attention from antitrust regulators. Same thing for the choice of Samsung 8nm for Ampere and TSMC 12nm for Turing, while AMD was on TSMC 7nm for both of those. Sure, volume was a large part of that decision, but they're already matching AMD with a 1-node deficit (Samsung 8nm is a 10+, and the gap between 10 and TSMC 7 is huge to begin with) and they were matching with a 1.5 node deficit during the Turing generation (12FFN is a TSMC 16+ node - that is almost 2 full nodes to TSMC 7nm). They cannot just make arbitrarily fast processors that dump on AMD, or regulators will get mad, so in that case they might as well optimize for cost and volume instead. If they had done a TSMC 7nm against RDNA1 they probably would be starting to get in that danger zone - I'm sure they were watching it carefully during the Maxwell era too. (the people who imagined some giant falling-out between TSMC are pretty funny in hindsight. (A) NVIDIA still had parts at TSMC anyway, and (B) TSMC obviously couldn't have provided the same volume as Samsung did, certainly not at the same price, and volume ended up being a godsend during the pandemic shortages and mining. Yeah, shortages sucked, but they could still have been worse if NVIDIA was on TSMC and shipping half or 2/3rds of their current volume.) Of course now we may see that dynamic flip with AMD moving to MCM products earlier, or maybe that won't be for another year or so yet rumors are suggesting monolithic midrange chips will be AMD's first product. Or perhaps "monolithic", being technically MCM but with cache dies/IO dies rather than multiple compute dies. But with RDNA3 AMD is potentially poised to push NVIDIA a little bit, rather than just the controlled opposition we've seen for the past few generations, hence NVIDIA reportedly moving to TSMC N5P and going quite large with a monolithic chip to compete.
- torrance 4y agoI’m not using Frontier, but I am using Setonix which is a large AMD cluster being rolled out in Australia. All of AMD’s teaching materials are about ROCm so this is very much how they’re expecting it to be used. The real pain for us is that there’s no decent consumer grade chips with ROCm compatibility for us to do development on. AMD have made it very clear they only care about the data centre hardware when it comes to ROCm, but I have no idea what kind of developer workflow they’re expecting there.
- uniqueuid 4y agoInteresting. So what is your workflow right now?
- torrance 4y agoDevelop against CUDA locally. Port my kernels to ROCm, and occupy a whole HPC node for debugging and performance tuning for a week. It’s terrible. Edit: I should say that their recommendation is to write the kernels in ‘hip’ which is supposed to be their cross device wrapper for both cuda or ROCm. I’m writing in Julia however so that’s not possible.
- claforte 4y agoThe AMD software stack has been behind for a long time but I feel like we're finally catching up. I heard that HIP (and hopefully the rest of ROCM) is now supported on the RX6800XT consumer GPU... maybe that could help? BTW my team at AMD has been using Julia for ML workloads for a while. We should get in touch - maybe some of the lessons we learn can be useful to you. My email is claforte. The domain I'm sure you can guess. ;-)
- vchuravy 4y agoIf you are using Julia I would recommend looking at AMDGPU.jl and (pluging my own project here) KernelAbstractions.jl
- claforte 4y agoBTW have you tried `KernelAbstractions.jl`? With it you can write code once that will run reasonably fast on AMD or NVIDIA GPUs or even on CPU. One of our engineers just started using it and is pleased with it - apparently the performance is nearly equivalent to native CUDA.jl or AMDGPU.jl, and the code is simpler.
- JonChesterfield 4y agoThe rocm stack is one of the toolchains deployed on Frontier. With determination, llvm upstream and rocm libraries can be manually assembled into a working toolchain too. It's not so much trickle down improvements as the same code.
- mastax 4y agoThese supercomputer contracts typically have a large amount dedicated to software support. I remember reading on AnandTech (?) that AMD was explicitly putting a bunch of engineers on ROCm for this project. It's one of the reason companies like these contracts so much.
- pinhead 4y agoSurprisingly, ROCm support has been getting a lot better over the very recent years. In my experience the pytorch support is essentially seamless between CUDA and ROCm. Also, I know some popular frameworks like DeepSpeed have announced support and benchmarks on it as well: https://cloudblogs.microsoft.com/opensource/2022/03/21/supporting-efficient-large-model-training-on-amd-instinct-gpus-with-deepspeed/ https://cloudblogs.microsoft.com/opensource/2022/03/21/suppo...
- eslaught 4y agoYes, DOE is very interested in DL. I don't work on this personally, but you can see an example e.g. here [1, 2]. You can see in the first link they're using Keras. I'm not up to date on all the details (again, don't work on this personally) but in general the project is commissioned to run on all of DOE's upcoming supercomputers, including Frontier. [1]: https://github.com/ECP-CANDLE/Benchmarks https://github.com/ECP-CANDLE/Benchmarks [2]: https://www.exascaleproject.org/research-project/candle/ https://www.exascaleproject.org/research-project/candle/
- belter 4y agoThank you to the authors for not calling it the fastest computer in the world :-) and instead, as they should, the most powerful. Clock speed is not the only factor of course, as instruction per cycle and cache sizes have an impact, but for a pure measure of speed, the fastest still is: - For practical use, and non overclocked, the EC12 at 5.5 Ghz: https://www.redbooks.ibm.com/redbooks/pdfs/sg248049.pdf https://www.redbooks.ibm.com/redbooks/pdfs/sg248049.pdf or - An AMD FX-8370 floating in Liquid Nitrogen at 8.7 Ghz: https://hwbot.org/benchmark/cpu_frequency/rankings#start=0#interval=20 https://hwbot.org/benchmark/cpu_frequency/rankings#start=0#i...
- formerly_proven 4y agoI guarantee you an FX-8370 isn't even close to being the fastest CPU even at 10 GHz. I bet most desktop CPUs you can buy nowadays will be faster out of the box.
- belter 4y agoTell me what your measure of fast is?
- PartiallyTyped 4y agoIt's embarrassing how slow that thing is compared to CPUs 2 years ago... The video below compares 8150 against CPUs from 2020 (i.e. no 5900x or 12900KS), includes data from 8370. https://youtu.be/RpcDF-qQHIo?t=425 https://youtu.be/RpcDF-qQHIo?t=425
- mihaic 4y agoThe more powerful processors become, the less I feel there's a need to build supercomputers. Thinking about it, the most powerful supercomputer in the world is pretty much a million consumer processors, working in parallel. That's going to stay pretty constant, since cost scales roughly linearly. If X is the processing power of $1k of consumer hardware, the bigger X gets, the less there is a difference in the class of problems that you can solve with X or X * 1e6 processing power.
- mastax 4y agoThe coherent memory interconnects between nodes is typically what makes supercomputers different than just a bunch of consumer hardware. It allows different types of programming or at least makes them easier.
- jabl 4y agoIt's a very fast, very low latency network fabric. But it's not coherent in the sense of cache coherent multiprocessors, and it doesn't offer shared memory style programming where you'd just load/store to addresses that happen to be mapped to another compute node somewhere in the system.
- l33t2328 4y agoI thought DMI allowed for exactly those kinds of load/store operations
- uniqueuid 4y agoSure, but consumer hardware does not have infiniband or other high-bandwidth interconnects. That means you can have at most ~1-2TB of ram accessible at any point. Some problems need coordination, and when you're back at OpenMP etc., a supercomputer suddenly makes sense.
- mihaic 4y agoI agree right now, I'm thinking maybe in 15 years you can have >1PB on a single machine, and then those problems that don't fit in that space but that fit in a supercomputer become fewer. 2050 will be within out lifetime. Basically I'm estimating the benefit ratio to be (log SupercomputerSize - log ConsumerSize)/log ConsumerSize, and that keeps decreasing.
- gsibble 4y agoWhat an incredible achievement. Good for AMD. The Epyc is a fantastic processor. And there are another 2 (3?) faster systems coming online in the next year or so.
- adrian_b 4y agoBesides being the first system exceeding the 1 Exaflop/s threshold, what is more impressive is that this is also the system with the highest ratio between computational speed and power consumption (i.e. the AMD devices have the first place in both Top500 and Green500). The AMD GPUs with the CDNA ISA have surpassed in energy efficiency both the NVIDIA A100 GPUs and the Fujitsu ARM with SVE CPUs, which had been the best previously. Unfortunately, AMD has stopped selling at retail such GPUs suitable for double-precision computations. Until 5 or 6 years ago, the AMD GPUs were neither the fastest nor the most energy-efficient, but they had by far the best performance per dollar of any devices that could be used for double-precision floating-point computations. However, when they have made the transition to RDNA, they have separated their gaming and datacenter GPUs. The former are useless for DP computations and the latter cannot be bought by individuals or small companies.
- visarga 4y agoComputational speed is important, but more important is the data transfer speed. At least in ML. Is AMD the best for data transfer speed?
- Const-me 4y ago> The former are useless for DP computations Looking at “double-precision GFlops” columns there [1] they don’t seem terribly bad, more than twice as fast compared to similar nVidia chips [2] While specialized extremely expensive GPUs from both vendors are way faster with many TFlops of FP64 compute throughput, I wouldn’t call high-end consumer GPUs useless for FP64 workloads. The compute speed is not terribly bad, and due to some architectural features (ridiculously high RAM bandwidth, RAM latency hiding by switching threads) in my experience they can still deliver a large win compared to CPUs of comparable prices, even in FP64 tasks. [1] https://en.wikipedia.org/wiki/Radeon_RX_6000_series#Desktop https://en.wikipedia.org/wiki/Radeon_RX_6000_series#Desktop [2] https://en.wikipedia.org/wiki/GeForce_30_series#GeForce_30_(30xx)_series_for_desktops https://en.wikipedia.org/wiki/GeForce_30_series#GeForce_30_(...
- scardycat 4y agoCongratulations to AMD, HPE and ORNL! This is an amazing achievement. Can't wait to see the spectacular science results coming from this installation. Intel was supposed to build the first Exascale system for ANL [1] [2]. to be installed by 2018. They completely and utterly messed up the execution, partly drive by 10nm failure, went back to the drawing board multiple times, and now Raja switched the whole thing to GPUs, a technology that Intel has no previous success with and rebased it to 2 ExaFlops peak, meaning they probably expect 1 EF sustained performance, a 50% efficiency. No other facility would ever consider Intel as a prime contractor again. ANL hitched their wagon to the wrong horse. 1. https://www.alcf.anl.gov/aurora https://www.alcf.anl.gov/aurora 2. https://insidehpc.com/2020/08/exascale-exasperation-why-doe-gave-intel-a-2nd-chance-can-nvidia-gpus-ride-to-auroras-rescue/ https://insidehpc.com/2020/08/exascale-exasperation-why-doe-...
- throwawaylinux 4y agoWhat is Raja?
- interesting_pt 4y agoI worked at Intel in a very closely related area. I quit after getting vaccinated for COVID, only stayed because of the pandemic. The biggest problem was that Intel simply couldn't execute. They couldn't design and manufacture hardware in a timely manner without too many bugs. I think this was due to poor management practices. My direct manager was amazing, but my skiplevel was always dealing with fires. It felt like instead of the effort being orchestrated that someone approached a crowd of engineers and used a bullhorn to tell them the big goal and that was it. The left hand had no idea what the right hand was doing. I often called Intel an 'ant hill', because the engineers would swarm a project just like ants do a meal. Some would get there and pull the project forward, some would get on top and uselessly pull upward, and more than I'd like would get behind the project and pull it backwards. Just a mindless swarm of effort, which generally inefficiently kinda did the right thing sometimes. The inability to execute started to effect my work. When I got a ticket to complete something, I just wouldn't. There was a very good chance that I'd have an extra few weeks (due to slippage) or the task would never need to get done, because the hardware would never appear. Planning was impossible. Conversely, sometimes hardware CAME OUT OF NOWHERE, not simple stuff, but stuff like laptops made by partners. Just randomly my manager would ask me to support a product we were told directly wouldn't exist, but now did. I needed to help our partner with support right now. Our partners were starting to hate us and it was palpable in meetings. I'm so glad I quit, I was being worked to the bone on a project which will probably fail and be a massive liability. Even if the economy crashes, and I can't get a job for years, and end up broke, it'll still have been worth it. I also only made 110K/yr base.
- photochemsyn 4y agoI wonder if having one supercomputer with x number of chips or having eight supercomputers each with x/8 number of chips would be the more practical working setup. Weather forecasting for example is basically a complex probabilistic algorithm, and there's a notion that running eight models in parallel and then comparing and contrasting the results will give better estimates of actual outcomes than running one model on a much more powerful machine. Is it feasible to run eight models on one supercomputer, or is that inefficient?
- derac 4y agoYou can run many programs on one supercomputer simultaneously, yes. Check out XSEDE. Cost-wise one big is going to be cheaper than 8 small due to infrastructure issues - cooling, maintenance, space, etc.
- pphysch 4y ago"XSEDE" proper is getting EOL'd in a couple months and transitioning to ACCESS [1]. [1] - https://www.hpcwire.com/off-the-wire/nsf-announces-upcoming-transition-from-xsede-to-access/ https://www.hpcwire.com/off-the-wire/nsf-announces-upcoming-...
- timbargo 4y agoYou can partition a large compute cluster into many smaller ones. Users can make a request specifying how many processors they want for how long. Check out this link to see the activity of a supercomputer at Argonne. https://status.alcf.anl.gov/theta/activity https://status.alcf.anl.gov/theta/activity And I believe it is more efficient to have a single large cluster. As there are large overheard costs of power, cooling, and having a physical space to put the machine in. Plus a personnel cost to maintain the machines.
- dang 4y agoThis reads more or less like a corporate press release - (edit: actually, it reads exactly like a corporate press release) - is there a more substantive article on the topic?
- eslaught 4y agoIt's not an article, but there's always the front page for the supercomputer (includes some limited specs): https://www.olcf.ornl.gov/frontier/ https://www.olcf.ornl.gov/frontier/ There's also detailed architecture specs on Crusher, an identical (but smaller) system: https://docs.olcf.ornl.gov/systems/crusher_quick_start_guide.html https://docs.olcf.ornl.gov/systems/crusher_quick_start_guide...
- mrb 4y agoI like this one, it gets into the specifics of the hardware, specifically the 7 slides in the middle of the article: https://www.tomshardware.com/news/amd-powered-frontier-supercomputer-breaks-the-exascale-barrier-now-fastest-in-the-world https://www.tomshardware.com/news/amd-powered-frontier-super...
- dang 4y agoOk, we changed to that from https://venturebeat.com/2022/05/30/amd-powers-worlds-most-powerful-supercomputer/ https://venturebeat.com/2022/05/30/amd-powers-worlds-most-po.... Thanks!
- briffle 4y agoWhat blows my mind is the newest NOAA super computer (that triples the speed of the last one) is a whopping 12 petaflops. It comes online this summer. It kind of shows the difference in priority spending, when nuclear labs get >1000 petaflop super computers, and the weather service (that helps with disasters that affect many Americans each year) gets a new one that is 1.2% of the speed. https://www.noaa.gov/media-release/us-to-triple-operational-weather-and-climate-supercomputing-capacity#:~:text=The%20computers%20%E2%80%94%20each%20with%20a,Venus%22%20in%20Orlando%2C%20Florida https://www.noaa.gov/media-release/us-to-triple-operational-....
- mulmen 4y agoWould a faster computer improve outcomes for victims of natural disaster? How much is left undiscovered about weather? Research spending is based on the potential for discovery. As a species we have studied weather since the beginning of time. How long have we been doing nuclear research? A century? Is there even an opportunity cost here? Or is it an economy of scale? As we build more supercomputers the costs go down. So NOAA and ORNL both get what they need for less.
- Twirrim 4y ago> Would a faster computer improve outcomes for victims of natural disaster? How much is left undiscovered about weather? The US is way behind on weather modelling, in part due to lack of computing power available to do the grids at sufficiently small cells compared to Europe and other parts of the world. That means less accurate predictions and less advance notice of impending disasters, which means more risk of loss of life and impact on infrastructure and the economy (and vice versa, inaccuracy can lead to more caution than is necessary, which has economic impact too). The US has to lean on Europe etc. for predictions. https://cliffmass.blogspot.com/2020/02/smartphone-weather-apps-can-you-trust.html https://cliffmass.blogspot.com/2020/02/smartphone-weather-ap... Talks about the fact that IBM / Weather.com actually uses a more accurate system than the NWS uses, because the NWS is still stuck on GFS (been several years now since congress passed an act to force NOAA to update away from it, and unfortunately it takes time)
- moffkalast 4y agoNow the real question: Can it run Crysis... without hardware acceleration?
- cesarb 4y ago> Can it run Crysis... without hardware acceleration? I understand you are joking, but it's a legitimate benchmark, one which I've seen at least Anandtech using. For instance, a quick web search found an article from last year (https://www.anandtech.com/show/16478/64-cores-of-rendering-madness-the-amd-threadripper-pro-3995wx-review/4 https://www.anandtech.com/show/16478/64-cores-of-rendering-m...) which shows an AMD CPU (a Ryzen 9) running Crysis without hardware acceleration at 1080p at nearly 20 FPS. As that article says, it's hard to go much higher than that, due to limitations of the Crysis engine.
- jamesredd 4y agoChina has two exaflop supercomputers. It's doubtful whether this is the world's most powerful supercomputer. https://www.nextplatform.com/2021/10/26/china-has-already-reached-exascale-on-two-separate-systems/ https://www.nextplatform.com/2021/10/26/china-has-already-re...
- ouid 4y agoI'm not really sure why you would trust this claim from China. Its not impossible, but its also not impossible to lie about
- sekia 4y ago> I'm not really sure why you would trust this claim from China. Why not? While I don't remember what was the previous US's x86 cluster that ranked as top of Top500 List (RoadRunner in 2009?), China's Tianhe-3 and OceanLight are direct successors of Tianhe-2A and TaihuLight, which are once fastest and still in top 10. These seems more promising to me.
- neo_blackcap 4y agoNYT claimed so https://www.nytimes.com/2022/05/30/business/us-supercomputer-frontier.html https://www.nytimes.com/2022/05/30/business/us-supercomputer...
- dekhn 4y agoall that matters in this context is whether they run TOP500 or not.
- curiousgal 4y agoHow much of that performance will get undone by the software though? Either through AMD's lack of effort or Intel's compiler "sabotage".
- ghc 4y agoIt probably won't be a factor. The likelihood of the system using standard compilers or drivers is quite low. It's non-trivial to optimize a compiler and drivers for a supercomputer, so companies like Cray make their own.
- robswc 4y agoSeems I've heard nothing but good things about AMD for the last 10 years or so. I once had an terrible experience with AMD ~10 years ago that made me swear off them for good. Had something to do with software but I remember it taking several days of work/solutions. Willing to give them another try soon though. I never seem to even use the full power of whatever CPU I get, lol.
- verst 4y agoLate 2020 I switched from Intel to AMD Ryzen 5900X for my gaming PC and only had great experiences as far as gaming is concerned. I should point out that there were significant USB problems on AMD B550, X570 chipsets (eventually addressed via BIOS updates). Unfortunately some professional audio gear is only certified for use with Intel chipsets and I have experienced some deal-breaking latency issues with ASIO drivers. For gaming I will be happy to continue using AMD - but for music I will probably switch back to Intel for my next rig.
- robswc 4y agoThat actually sucks because I do a lot of music stuff and having any issues with ASIO would be a deal breaker. Thanks for the heads up! One of those things I would have never even thought of to check! Also sums up my AMD experience 10 years ago. Stuff just wasn't working :/
- UberFly 4y agoLisa Su joined AMD in 2012 and in 2017 the first Zen chips were released. Good people making good decisions.
- zepmck 4y agoThe most powerful and unfortunately unusable supercomputer of the world. AMD's approach to GPUs is on a failing track since its inception. The only software stack available is super fragile, buggy and barely supported. Rather than building a HPL machine I would have preferred see public money spent in a different way.
- ghc 4y agoIt's a supercomputer. The programming model is very, very different. The software stack is full of incredibly fragile stuff from any number of manufacturers. It's honestly hard to even describe how much more difficult using MPI with Fortran on a supercomputer is compared to anything I've ever touched elsewhere. Maybe factory automation comes close?
- wait_a_minute 4y agoHow could someone get practical experience in this space?
- ghc 4y agoI know of five ways: 1. As an undergraduate, join a research group that needs to run simulations on a supercomputer. 2. As a grad student, join a research group that works with supercomputers. 3. As a software engineer or IT person, join a research group at a university. They need people too, but fair warning: the pay is...subpar. 4. Join a national laboratory in some capacity. This route necessitates working for your country's government or military, which may or may not be palatable to you depending on how you feel about your gov't/military. 5. Join a giant multinational company that has supercomputers and uses them. Exxon is a good example. They have massive supercomputing power. Unless you're an undergrad, I'm afraid all the ways I know of suck in some way or another. I did 1 & 3. As for the rest, I think 2 would make the most sense if you have BS, because you can go get a masters in a year or so while getting the experience.
- mrb 4y agoThe HN crowd would probably prefer reading the many technical details at the ORNL press release: https://www.ornl.gov/news/frontier-supercomputer-debuts-worlds-fastest-breaking-exascale-barrier https://www.ornl.gov/news/frontier-supercomputer-debuts-worl... which I just submitted here: https://news.ycombinator.com/item?id=31573066 https://news.ycombinator.com/item?id=31573066 Also, yesterday Tom's hardware had a detailed article: https://www.tomshardware.com/news/amd-powered-frontier-supercomputer-breaks-the-exascale-barrier-now-fastest-in-the-world https://www.tomshardware.com/news/amd-powered-frontier-super... 29 MW total, 400 kW per rack(!) And, anyone else is like me and wants to see actual pictures or videos of the supercomputer, instead of a rendering like in venturebeat article? Well, head here, ORNL has a very short video: https://www.youtube.com/watch?v=etVzy1z_Ptg https://www.youtube.com/watch?v=etVzy1z_Ptg We can see among other things: that it's water-cooled (the blue and red tubing), at 0m3s we see a PCB labelled "Cray Inc Proprietary ... Sawtooth NIC Mezzanine Card"
- pvg 4y agoNot much point submitting a dupe with the discussion already on the front page but you can email your better links to the mods who are looking for better a better link: https://news.ycombinator.com/item?id=31571551 https://news.ycombinator.com/item?id=31571551
- SoftTalker 4y agoSince Cray stopped making their own CPUs, they have been back and forth between AMD and Intel several times.
- wmf 4y agoIt's not really back and forth; Cray supports Intel, AMD, and ARM CPUs equally as well as Nvidia, AMD, and Intel GPUs.
- marcusjramsey 4y agohmm
- anuvrat1 4y agoCan someone please explain, how software is made at this scale?
- sydthrowaway 4y agoUsing a HPC Framework, such as OpenMP
- mhh__ 4y agoFairly low tech until you get to the super high end. You have a blend of very specific domain specific knowledge (e.g. they know the hardware - the interconnects more than the CPUs) and old skool Unix system administration.
- peter303 4y agoOne petaflop DP linpack achieved in 2008. Supercomputing "Moores Law" is doubling speed every 1.5 years, order of magnitude every five years, a thousand-fold 15 years. Pretty close to schedule. Onward to a zettaflop around 2037?
- gigatexal 4y agoI am still kicking myself every time I look at AMD’s share price. I sold a not-insignificant-to-me amount of shares when the price was basically below 10 a share. Now it’s above 100. All this is to say that the turn around at AMD is good to see and the missteps at Intel are hilarious. This is like the time the Athlon64 and it’s on die memory controller was kicking the Pentiums around.
- BooneJS 4y agoWhile AMD gets top billing for the compute cores, HPE used the acquired Cray Slingshot network to create this heterogeneous supercomputer. It has a 64-port, 12.8 Tb/s bandwidth switch, it scales to >250,000 host ports with maximum of 3 hops, and it uses Ethernet "plus optimized HPC functionality".
- vfclists 4y agoI read somewhere that this means the US now has the world's fastest supercomputer. Does this No. 1 position have something to do with the ban on exporting advanced technology to China?