6 ms·
Why Fugaku, Japan's fastest supercomputer, went virtual on AWS
- echelon 2y agoI can't even imagine the hourly bill. This seems like a great way for Amazon to gobble up institutional research budgets.
- KaiserPro 2y ago> I can't even imagine the hourly bill Yes, and no. As a replacement for a reasonably well used cluster, its going to be super expensive. Something like 3-5x the price. For providing burst capacity, its significantly cheaper/faster, and can be billed directly. But would I recommend running your stuff directly on AWS for compute heavy batch processing? no. if you are thinking about hiring machines in a datacenter, then yes.
- littlestymaar 2y ago> For providing burst capacity, its significantly cheaper/faster, and can be billed directly. Cloud is cheaper for such a workload, yes. But you still wouldn't want to pay the AWS premium for that. But I guess nobody ever got fired for choosing AWS.
- KaiserPro 2y agoDepends. If you want high speed storage, thats reliable and zero cost to setup (ie doesn't require staff to do) then AWS and spot instances is very much a cheap option. If you have in house expertise and can knock up a reliable high speed storage cluster (no porting to s3 is not the right option, that takes years and means you can't edit files inline.) then running on one of the other cloud providers is an option. But they often can't cope with the scale, or don't provide the right support.
- littlestymaar 2y ago> and zero cost to setup (ie doesn't require staff to do) Ah, the fallacy of the cloud not requiring staff to magically work. Funnily enough, every time I heard this IRL, it was coming from dudes who was in fact paid to manage AWS instances and services for their company as their day job but somehow never included his salary and the one of his colleagues. > But they often can't cope with the scale Most people vastly overestimate the scale they need or underestimate the scale of even second tier cloud providers but there's not many workload that could cause storage trouble anywhere. For the record a single rack can host 5/10 PB, how many people needs multiple exabytes of storage again? > or don't provide the right support. I've never been a whole-data-centers-scale customer but my experience with AWS support didn't leave me in awe.
- KaiserPro 2y ago> Ah, the fallacy of the cloud not requiring staff to magically work. its not a fallacy, its a tradeoff. If you want to have 100nodes doing batch work in a datacentre, you need to buy, install, commission and monitor hardware, storage and networking. Now, That is possible, but hard to do on your own. So realistically you'd get a managed service to install that stuff in a data centre. As your experience will show you, the level of service you get is "patchy". You really need some one who has worked on real steel to specify that kind of work. Those people are incredibly rare, especially if you want someone who can talk the rest of the stack as well. (That is one massive downside of the cloud, there are hardly any new versatile sysadmins being created.) As a former VFX sysadmin who looked after a number of large HPC clusters, its fairly simple to get something working, but something running reliably and fast is another matter. > Most people vastly overestimate the scale they need Yes, but we are talking about an HPC cluster that is >100k nodes. Something fast enough to cause hilarious QoS issues on most storage systems. especially single rack arrays with undersized metadata controllers. Even on isilon with its "scale out" clustering, one job can easily nail the entire cluster by doing too many metadata operations. (protip single namespace clusters are rarely worth the effort outside of a strict set of problems. Its much better to have a bunch of smaller fileservers grouped together in an automount. Think of it like DB sharding, but with folders.) Second tier cloud providers are rarely able to provide 10k GPUs on the spot, certainly not at a reasonable price. Most don't have the concept of spot pricing, so you're deep into contract negotiations. > AWS support didn't leave me in awe. Its not great unless you pay. but you will eventually get to someone who can answer your question. Azure on the other hand, less so.
- moralestapia 2y agoI currently work on a research group that does all of its bioinformatics analysis in AWS. The bill is expensive, sure, 6 figures regularly, maybe 7 in its lifetime, but provisioning our own cluster (let alone a supercomputer) would've been way more costly and time consuming. You also now need people to maintain it, etc... And at the end of the day, it could well be the case that you still end up using AWS for something (hosting, storage, w/e) I think it's a W, so far.
- deleted 2y ago[deleted]
- fock 2y agobecause we essentially did spend exactly that to build a cluster: what amount of ressources are you using (coreh/instances/storage)?
- jasonjmcghee 2y agoThere are tradeoffs in both cases. Some issues with building a cluster are, you're locked into the tech you bought, need space, expertise to manage it, cost to run/maintain. You can potentially recoup some of that investment by selling it later, usually a quickly depreciating asset though. But, no AWS or AWS premium.
- littlestymaar 2y agoThere's also the third way, of going for intermediate-size cloud providers, you get the lack of Capex and actual hardware to deal with without the AWS premium. I don't understand why so many people act as if the only alternative was AWS or self-hosted.
- moralestapia 2y agoNever tried to suggest that, but tbh, the value added services on both AWS and GCP are hard to emulate, and they "just work" already. Sure, I could spend a few weeks (months?) compiling and trying some random GNU-certified "alternatives" for Cloud Services but ... just nah ...
- prmoustache 2y agoI guess it really depends how many hours per year the Fugaku is used at its max capacity. Also in this case, it could grow progressively.
- Something1234 2y agoSomeone needs to share utilization data on these super computer clusters. Most have a long queue of jobs waiting to be ran, and you have to specify how long you think your job is gonna take to properly schedule it.
- exe34 2y agoAt this scale, I would expect they have better terms than the retail pricing.
- edu 2y agoThe title is a little bit misleading, Fugaku stills exists phisically. This project is about being able to replicate the software environment used in Fugaku on AWS.
- sheepscreek 2y agoThey haven’t said a word about costs. Does it cost as much/less/more? One challenge would be deciding on the unit to compare the costs. They could pick cost per Peta/ExaFLOPs.
- throw0101b 2y agoIf anyone wants to dabble with (HPC) compute clusters, ElastiCluster is a handy tool to spin up nodes using various cloud APIs: * https://elasticluster.github.io/elasticluster/ https://elasticluster.github.io/elasticluster/
- aconz2 2y agoThe goal seems to be: "Virtual Fugaku provides a way to run software developed for the supercomputer, unchanged, on the AWS Cloud". Is AWS running A64FX's? Is there more info on all the details of this somewhere? I think it is a compelling goal to have portability across compute and am curious how they are going about it and what tradeoffs they've made. Software portability is one level of hard and performance portability is another. I wish we could get an OCI spec for running a job that could fill this purpose
- tame3902 2y agoThis article has some more details: https://www.hpcwire.com/2023/01/26/riken-deploys-virtual-fugaku-on-aws/ https://www.hpcwire.com/2023/01/26/riken-deploys-virtual-fug... It looks like they want to make the transition between Fugaku and AWS and vice versa easier.
- glial 2y agoThe article doesn't mention this, but I imagine having an AWS version of the supercomputer would be extremely helpful for software development and performance optimization, especially if the same code could be tested with fewer nodes.
- adev_ 2y agoThere is a big interest of what Fugaku and Dr. Matsuoka are doing here and it seems that this article is missing it entirely. HPC development is not your standard dev workflow where your software can be easily developed and tested locally on your laptop. Most software will requires a proper MPI environment with a parallel file system and (often) a beefy GPU. Most development on a supercomputer is done on a debug partition. A small subset of the supercomputer reserved for interactive usage. That allows to test the scalability of the program, hunts Heisenbug related to concurrency, access large datasets, etc... But Debug partitions are problematic: Make it too small and your scientists & devs loose productivity. Make it too large and you are wasting your supercomputer resources to something else than production jobs. The Cloud solves this issue. You can spawn your cluster, debug and test your jobs, destroy your cluster. You do not need very large scale nor extreme performance, you need flexibility, isolation and interactivity. The Cloud gives you that because of the virtualization.
- bch 2y ago> Most software will requires a proper MPI environment with a parallel file system and (often) a beefy GPU I’m but a tourist in this domain, but can you dig into this a bit more and compare/contrast w “traditional” development? I presume the MPI you’re talking about is OpenMPI or MPICH, which need to be dealt with directly - but what are the considerations/requirements for a parallel FS? Hardware is hardware, and I guess what you’re saying re: GPUs is that you can’t fake The Real Thing (fair enough), but what other interesting “quirks” do you run into in this HPC env vs non-HPC?
- avidphantasm 2y agoLots of legacy HPC code assumes POSIX file I/O, which means a parallel file system, which means getting the interconnect topology right, which is not easy.
- trueismywork 2y agoThe almost single most important feature of supercomputers is guaranteed low latency interconn3ct between two nodes of a supercomputer, which guarantees high performance for even very talkative workloads. This is why supercomputers document their network topology so thoroughly and allow people to reserve nodes in a single rack for example.
- astrodust 2y agoToday "supercomputer" translates to "really high AWS spend".
- hn_throwaway_99 2y agoAs someone not versed in the field, can anyone explain the types of workloads that are ideal for HPC machines, as opposed to, for example, a huge number of networked GPUs? E.g. it reportedly cost over $100 million to train GPT 4. My understanding is this was done in a huge number of high performance GPUs, e.g. A100s, etc. So I guess my question then is what would a dedicated "supercomputer" be used for today that couldn't be accomplished on one a more traditional network? To emphasize, none of these questions are meant to be rhetorical or leading - I honestly don't know and am curious.
- bbatha 2y agoThat’s basically what super computers have been for the last twenty years: large computer optimized data centers. The main differences in hardware are usually infiniband (and its RDMA capabilities) paired with a really powerful parallel file system cluster. Occasionally they’ll be exotic compute accelerators or lately just variants of the gpus specific to the cluster. On the software side it’s usually slurm managing code that leverages MPI, openMP and CUDA. Having a large homogeneous cluster that has a ~10 year lifespan means that you you have a lot of tuning up and down the stack from specific MPI implementations for your infiniband hardware to optimization tricks in the science code tuned for the specific gpu and cpu models. All of this has gotten even less specialized since GPUs started to take off and those trends have accelerated with ML having the same needs. ML is also prompting clouds to provide the same kinds of hard ware with infiniband and rdma.
- trueismywork 2y agoAlgorithms like differential equation solvers for extremely fast wave speed physical systems. Molecular dynamics, electronic structure calculations. Any algorithm that requires FFT over 1 billion cells.