3 ms·
A lot of MPI’s (ab)use in HPC boils down to distributed task management in lieu of a work queue system available to users. People have embarrassingly parallel j
by prpl 4y ago
A lot of MPI’s (ab)use in HPC boils down to distributed task management in lieu of a work queue system available to users. People have embarrassingly parallel jobs but need to coordinate on the task management because many HPC centers either don’t provide resources for a long-lived service to execute near the cluster (or even general connectivity outside the cluster)
The problem is that you do have to support true parallel MPI jobs in those shared clusters though, so MPI just becomes the hammer for everyone else.
Managing the resources a level higher (all resources live in a k8s cluster, slurm under k8s) seems to be the best way to really accommodate both types of loads but most HPC centers are far off from implementing that.
https://slurm.schedmd.com/SC22/Slurm-and-or-vs-Kubernetes.pdf https://slurm.schedmd.com/SC22/Slurm-and-or-vs-Kubernetes.pd...
(I think that presentation has some misconceptions about k8s - most k8s clusters are elastic to a max size - and it sounds like the really want to control most scheduling - but it gives an overview of merging the systems)
- saltcured 4y agoRoughly 20 years ago, the Condor high-throughput computing system gained "glide-ins" to do this sort of repurposing. Before that, Condor was mostly persistent runners on desktop fleets. After, you would submit a batch job to an HPC cluster and for the duration of the job, those HPC nodes became additional runners for an existing Condor scheduler. Around that time, there was also a period of reservation-based "advanced scheduling" where HPC centers were flirting with making future scheduling promises. The right way to think of these would be like guarantees to get bare metal machine capacity during a certain wall-clock period. In my opinion, the commercial pre-cloud/cloud/virtualization stuff then infected everyone and regressed to time-sharing with fuzzy QoS and lots of over-subscription and dynamic rescheduling. Of course, these different approaches will all be isomorphic in the end if they explore the full space of application requirements. The traditional paths were just approaching from very different economic priorities. The IaaS folks are incrementally adding more QoS and pricing options which could eventually provide HPC IaaS if carried to full fruition. I.e. future guarantees of significant hardware resources. But as far as I know, those are still in the realm of "talk to a sales rep" and not some automated IaaS request flow at this point.
- dekhn 4y agoDuring the years I was active in grid computing (before cloud computing became huge) Miron Livny (creator of Condor) would basically attend every talk and explain how "condor already does this, why are you reinventing the wheel?"
- saltcured 4y agoHah, yes. And Oracle reps used to say the grid was inside their cluster. A lot of folks did not appreciate the decentralization of the grid. The grid computing concepts were about federation of disparate organizations and their resources, not about how sprawling of a system a single vendor or HPC center could build.
- tannhaeuser 4y agoI believe the terminology is/was "advance reservations" (as in reservations of resources such as CPU, mem, disk space, i/o and net bandwith in advance on clusters otherwise freely available to ad-hoc jobs) rather than "advanced" anything, or at least it was with the Torque scheduler I reviewed for a clickstream analysis customer project.
- saltcured 4y agoYes, I think so too. I blame the predictive typing in my hands.
- prpl 4y agoThe PITA part of condor/grid was software management before containers. Sure, everyone was running at least RHEL4/5/6 (Or SL4/5/6) and in many cases AFS worked and the more advanced operators were adding VM execution, but it was (and still is) annoying to deal with. Most annoying right now is that nobody can agree on a container runtime - there’s Docker/Shifter, singularity, Charlie. It should just all be podman now but everybody is still holding on. (I have worked with Slurm, condor, Torque/PBS, gridEngine, DIRAC, and LSF) Plugging a different scheduler into k8s might be an interesting way of solving this - it seems like there’s a lot of work on scheduler plugins when I last looked. Some of the issues are similar in cloud too - coscheduling by latency. There’s at least some incentive to not solve this - I remember k8s on Mesos being popular and of course we know how that played out.