5 ms·
AMD's Instinct MI455X: Aiming for the Sun
- thyristan 3mo ago> The basic GCN microarchitecture underpinned every single one of AMD’s compute accelerators for nearly 15 years [...]. With CDNA5 AMD has moved over to a microarchitecture that is based on the RDNA series putting a bookend to the long-lived line that was the GCN microarchitecture. So everything is new again. Will this be another "buy now, have ROCM work in 3 maybe years if you are lucky"?
- p_l 3mo agoArguably this might reduce it, because now ROCm won't be prioritised for the GCN-based systems?
- Farfignoggen 3mo agoAt least going forward, with both AMD's AI accelerators and Consumer GPUs using the same RDNA based IP, there will be less reason for AMD to drop RDNA GPUs from the ROCm/HIP support matrix! And so with both MI series accelerators and consumer GPUs/Graphics all using the same basic RDNA Micro-Architecture the ROCm/HIP code base developed for "CDNA5"/later will work for consumer RDNA as well with minor changes required! Polaris and Vega GPUs/Graphics both have been dropped from the ROCm/HIP support matrix. But with AMD's "The Rock" software stack making use of SPIR-V in the same manner as Nvidia/PTX there made be little issues getting that to work with older Vega/Polaris GPUs. even if AMD's not validated/verified that software stack with Consumer Vega/Earlier Consumer GPUs and Integrated Graphics. AMD's biggest issue has been dropping its older GPU micro-architectures from the ROCm/HIP support matrix too soon, and just look at Blender 3D's HIP Back End that requires RDNA2/Later GPUs and Graphics for any Radeon iGPU/dGPU Accelerated Blender 3D Cycles rendering support!
- nolist_policy 3mo agoNot anymore: https://rocm.blogs.amd.com/software-tools-optimization/spir-v-rocm/README.html https://rocm.blogs.amd.com/software-tools-optimization/spir-...
- dist-epoch 3mo ago432 GB of RAM per chip. Tens of terabytes per rack. Clients queuing up to buy them. RAM prices are not coming down any time soon.
- nl 3mo agoDid anyone think RAM prices were coming down soon?
- ACCount37 3mo agoOf course. Wishful thinking is a true staple of human intelligence.
- Goronmon 3mo agoThere are plenty of laymen that think "AI bubble pop" means that "AI" usage will go to 0 and then all of a sudden all these hardware and datacenters related companies are buying will become totally unused and flood the market.
- swiftcoder 3mo agoHey, they are only projected to cost about $5.5million a rack. Pocket change, really
- stevefan1999 3mo agobut you have to rewrite all your software to ROCm. And ROCm, to this day, still sucks. ZLUDA basically crash on high memory demand, and only accounted for ~70% of CUDA API coverage (and it is still buggy). But hey, at least it does run on Rust-CUDA. I'm one of the few who ported it and fixed a few bugs on ZLUDA. I used it to run a simple SHA256 kernel and it ran sure, but I gave it up because of those fundamental problems on AMD GPUs. You can't believe how messy ROCm is. I wonder if Vulkan compute kernel using SPIR-V would be a better choice.
- dist-epoch 3mo agoCo-founder and Chief Compute Officer Anthropic: > I think the thing that we were thinking about originally was whenever we're bringing up a new hardware platform, it's a big effort. It's like a huge thing. And so as we were thinking about this, we started doing our own evaluation of MI 355. You guys generously got us a rack to start working. And we expected this to be kind of a big process. > Our actual experience was we had one engineer who start doing it. They spun up Claude, asked it, hey, bring up this machine, left it going over the weekend. And we ended up with a graph of the actual performance of our leading model on it, just going up and up and up over the weekend. https://x.com/austinsemis/status/2080336781782753635 https://x.com/austinsemis/status/2080336781782753635
- fancyfredbot 3mo agoIt's mind blowing seeing these multi exaflop single rack systems. The world's first exaflop supercomputer was Frontier. It was launched only 4 years ago in 2022. It's not a fair comparison of course. FP4 in Helios barely qualifies as floating point. Frontier was proper fp64, 16 times the bit width and probably 256x as many transistors. All the same just wow. Much compute.
- ACCount37 3mo agoWorkloads did change over time. Back when we were first approaching practical exascale, the dominant workload for a supercomputer was thought to be physics simulations - and they often benefit from high numerical precision. Now, the dominant compute-hungry workload is AI, where precision takes second place to the independent parameter count. To the point that the capacity of BF16, which were originally designed as a radical optimization for AI workloads, is sometimes considered wasteful now. AI workloads have some truly peculiar and counterintuitive properties - the kind of things you might expect to see in biology instead of conventional computing. Intrinsic error tolerance, for one. It did necessitate some rethinking and reprioritization, and I'm not quite sure if we converged to the general shape of an "optimal" AI accelerator as of yet.
- ux266478 3mo ago> the dominant workload for a supercomputer was thought to be physics simulations - and they often benefit from high numerical precision. It still is. Just like how "mainframe" used to be a very general word, and over time gained a very unintuitive definition referring to a very specific type of computer with a specific purpose, "supercomputer" almost invariably means it's a highly bespoke cluster dealing with FP64 workloads. I don't see anyone referring to these AI clusters as supercomputers, for the same reason they aren't referring to the racks as mainframes. I wonder if there's a term for this kind of semantic narrowing?
- fancyfredbot 3mo agoI think the term in this case is marketing. NVIDIA used to use the term supercomputer more often five or six years ago. It's now fully committed to the term AI factory. I think this is supposed to make you think of their expensive kit as a productive asset building your business rather than an expensive tool for your boffins. Your average business executive is instinctively going to question whether they need a supercomputer but don't yet have the same preconceptions about AI factories. I don't think that Nvidia single handedly caused the move away from the term supercomputer. But I think they had a hand helping push firmly in that direction.
- pixelpoet 3mo agoI guess I submitted the link at the wrong time or something.
- jingpostmedia 3mo ago[flagged]
- neogodless 3mo agoDuplicate post: https://news.ycombinator.com/item?id=49025398 https://news.ycombinator.com/item?id=49025398 (13 hours earlier)
- inigyou 3mo agoAn apt title. To get to the sun, you need to slow down from orbital speed to zero. And then you fall in. Do they run CUDA yet or is this yet more impressive but unusable hardware? AMD never misses a chance to miss a chance.
- lostmsu 3mo agoI don't have that many kidneys