4 ms·
That’s basically what super computers have been for the last twenty years: large computer optimized data centers. The main differences in hardware are usually i
by bbatha 2y ago
That’s basically what super computers have been for the last twenty years: large computer optimized data centers. The main differences in hardware are usually infiniband (and its RDMA capabilities) paired with a really powerful parallel file system cluster. Occasionally they’ll be exotic compute accelerators or lately just variants of the gpus specific to the cluster. On the software side it’s usually slurm managing code that leverages MPI, openMP and CUDA. Having a large homogeneous cluster that has a ~10 year lifespan means that you you have a lot of tuning up and down the stack from specific MPI implementations for your infiniband hardware to optimization tricks in the science code tuned for the specific gpu and cpu models. All of this has gotten even less specialized since GPUs started to take off and those trends have accelerated with ML having the same needs. ML is also prompting clouds to provide the same kinds of hard ware with infiniband and rdma.