4 ms·
There is a big interest of what Fugaku and Dr. Matsuoka are doing here and it seems that this article is missing it entirely. HPC development is not your stan
by adev_ 2y ago
There is a big interest of what Fugaku and Dr. Matsuoka are doing here and it seems that this article is missing it entirely.
HPC development is not your standard dev workflow where your software can be easily developed and tested locally on your laptop.
Most software will requires a proper MPI environment with a parallel file system and (often) a beefy GPU.
Most development on a supercomputer is done on a debug partition. A small subset of the supercomputer reserved for interactive usage. That allows to test the scalability of the program, hunts Heisenbug related to concurrency, access large datasets, etc...
But Debug partitions are problematic: Make it too small and your scientists & devs loose productivity. Make it too large and you are wasting your supercomputer resources to something else than production jobs.
The Cloud solves this issue. You can spawn your cluster, debug and test your jobs, destroy your cluster. You do not need very large scale nor extreme performance, you need flexibility, isolation and interactivity. The Cloud gives you that because of the virtualization.
- bch 2y ago> Most software will requires a proper MPI environment with a parallel file system and (often) a beefy GPU I’m but a tourist in this domain, but can you dig into this a bit more and compare/contrast w “traditional” development? I presume the MPI you’re talking about is OpenMPI or MPICH, which need to be dealt with directly - but what are the considerations/requirements for a parallel FS? Hardware is hardware, and I guess what you’re saying re: GPUs is that you can’t fake The Real Thing (fair enough), but what other interesting “quirks” do you run into in this HPC env vs non-HPC?
- avidphantasm 2y agoLots of legacy HPC code assumes POSIX file I/O, which means a parallel file system, which means getting the interconnect topology right, which is not easy.
- trueismywork 2y agoThe almost single most important feature of supercomputers is guaranteed low latency interconn3ct between two nodes of a supercomputer, which guarantees high performance for even very talkative workloads. This is why supercomputers document their network topology so thoroughly and allow people to reserve nodes in a single rack for example.
- gyrovagueGeist 2y agoThe interconnect and network topology is also a big component of the hardware where you can't "fake The Real Thing" in practice. You can often get fairly confident in program correctness for toy problem runs by scaling 1-~40 ranks on your local machine, but you can't tell much about the performance until you start running on a real distributed system where you can see how much your communication pattern stresses the cluster. Or if you run into bugs / crashes that needs 1000s of processes or a full scale problem instance to reproduce, god help you and your SLURM queue times.
- scheme271 2y agoDepends on the software being used but it's probably OpenMPI (or a variant). However, OpenMP is also used especially in a hybrid mode where on a node OpenMP is used for shared memory parallelism and OpenMPI is used for inter-node parallelism. The parallel FS stuff is mainly to handle large number of nodes streaming data in and out. E.g. a few thousand nodes all reading from data sets or saving checkpoint data. One big difference you'll see in HPC is large scale, fine grained parallelism. E.g. a lot of nodes simulating some process where you need to resync all the nodes and exchange data between them at each time step. Also checkpointing, i.e. since simulations may take weeks to run, most apps support saving application state to disk periodically so that if something crashes, you'll only lose a few hours or a day of computation. The checkpointing also causes a bunch of FS io since you need to save the application state from all the nodes to storage periodically so you'll see really high io spikes when that happens.