6 ms·
THP are a quick win for people who have little control over how malloc behaves. Like Python and numpy, for instance. I did a lot of HPC modelling ranging from
by mickeyp 4y ago
THP are a quick win for people who have little control over how malloc behaves. Like Python and numpy, for instance.
I did a lot of HPC modelling ranging from hundreds of GiB to TiB-sized RAM servers and THP was an instantaneous win over not using it. Later on I experimented with LD_PRELOAD_PATH and libhugetlbfs and while it did stabilise things even more and reduced time spent in the page table, it was not even several factors 'better' than THP.
A large part of this really boils down to your performance needs. Deterministic modelling -> fairly stable and reproducible malloc patterns -> THP will probably work OK.
If your memory usage spikes a lot then THP is probably more hassle than it's worth. The kernel will spend an eternity thrashing and kicking the page table and trying in vain to clean up after itself. I can see why people feel THP sucks under those conditions!
If anything, the real crime here is how awful Linux is at HPC without a lot of tuning and careful tweaking. Throw NUMA into the mix, file caches and the OOMkiller and it feels like it never moved out of the 90s. Combine it with poorly-configured hypervisors and your performance will yo-yo and you'll spend an eternity trying to figure out why (ask me how I know...)
- basilgohar 4y agoWhen you say: the real crime here is how awful Linux is at HPC without a lot of tuning and careful tweaking To what are you making the comparison? Are *BSDs better? HP-Unix? AIX? Windows? What's the competition in this space that makes Linux look bad in HPC? Edit: Added "in HPC" at the end for clarification.
- jsjohnst 4y agoI’m curious if GP has any positive examples too, but did want to mention that being “the best of the awful” doesn’t make one not “awful”.
- pixl97 4y agoTypically being 'best of awful' means you have both a difficult problem and a very small dataset to work with where the solutions are not easily generalizable. You have to look at each and every HPC workload as a custom application where very small changes in your input and environment could have massive behavior changes in the application performance.
- jsjohnst 4y agoI don’t disagree with you, but that doesn’t negate GP’s or mine’s point.
- trelane 4y agoI don't think there has to be a better, existing alternative to say something is terrible. AFAICT, this is just generally a hard space to get right, especially by default, on a system that will power both a data center - size supercomputer and system that is barely me then a microcontroller.
- pjmlp 4y agoThe competition is naturally UNIX based OSes, given the historical background of such systems. Even the workloads that use GNU/Linux, most likely are heavily customized versions provided by IBM, HP and co. https://www.ibm.com/high-performance-computing https://www.ibm.com/high-performance-computing https://www.hpe.com/us/en/compute/hpc.html https://www.hpe.com/us/en/compute/hpc.html I also remember that for a while IBM's xl compilers were used quite often, not sure if that is still the case. See https://www.top500.org/ https://www.top500.org/
- gnufx 4y agoThe competition probably shouldn't be full Unix, and in Blue Gene, for instance, it wasn't on the compute nodes. I don't know exactly what the current Cray environment is like -- it used to be rather odd -- but most HPC systems run normal EL-ish distributions. Summit is RHEL8, and mostly uses GCC, not XL according to https://gcc.gnu.org/wiki/cauldron2022talks?action=AttachFile&do=view&target=OpenMP-OpenACC-Offload-Cauldron2022-1.pdf https://gcc.gnu.org/wiki/cauldron2022talks?action=AttachFile...
- pjmlp 4y agoSure, however they are mostly UNIX inspired if you will, as anything else.
- gnufx 4y agoSingle-user, single-program doesn't seem very Unix-y to me: https://en.wikipedia.org/wiki/CNK_operating_system https://en.wikipedia.org/wiki/CNK_operating_system There was also Plan 9, but I don't know if it was ever actually used: https://doc.cat-v.org/plan_9/blue_gene/ https://doc.cat-v.org/plan_9/blue_gene/
- convolvatron 4y agoThere has been an ongoing thread that claims that hpc workloads need no kernel or very little and that sharing resources doesn’t make sense at that grain size. This makes sense to me
- gnufx 4y agoI agree about huge pages in HPC, but regarding Linux -- if you want high performance (at least on a system designed for differing workloads), surely you should expect to do performance engineering. Whether or not you should be doing HPC on such a kernel is another question, but alternatives haven't caught on, perhaps because of the way applications which they need to support are written. If OOMkiller kicks in it suggests a resource manager issue (not accounting memory). Then, you can't blame the kernel for NUMA, which you presumably want for performance; I don't understand the difficulty people have with pinning processes after many years of it being necessary. There's something to be said for userspace filesystems too, like the venerable PVFS. The answer to hypervisors should be don't do that.
- mickeyp 4y agoI don't mind doing the engineering. But penetrating the veil of some of these things, ill-defined as they are, with work on certain areas being sporadic and taking place of years or decades, makes it especially hard. In fact it's the _absence_ of choice that hurts more than it is a surplus of choice. As for OOMkiller: I think we're both wise enough to know that experimentation is a large part of what any team that consumes gobs of RAM do. So talk of a "resource manager issue" is all well and good in prod when you should have a reasonable handle on what's used by what and for how long. Less so when you're scaling a model --- say a monte carlo model -- to a larger number of simulations in development and you're testing things. When you're paying an awful lot of money for hardware you start to count (or you should, if you respect your company's money) the costs of things like this. Regardless, the OOMKiller will indeed reap stuff; and more often than not, it'll pick something shouldn't (it's probabilistic and wrong as much as it is right) and that can cause headaches. As for NUMA: I'm not blaming the kernel for anything :-) And hypervisors are sometimes a given, and not a choice. We don't all get to pick our hardware, nor what hardware is made available to us.
- gnufx 4y agoDefinitely there's a lack of shared wisdom about tuning parameters (along with other things in HPC, at least in the circles I see), and too much churn with Linux v. userland (e.g. cgroups). Actually, I was forgetting the stupidity of the memory cgroup invoking the OOMkiller, rather than giving ENOMEM, if you purely use the cgroup for memory accounting -- a sensible-looking change from OpenVZ was rejected. That means there's probably no indication about what's happened unless the job is correlated with the syslog message. As I don't get to do that stuff any more, I haven't looked for a way to hook it now. There's obviously a window in which it can fail, but the resource manager can track the job's PSS, as well as at least ulimiting data and stack. You certainly don't want the job to start paging, if you have swap -- that definitely wastes the resource. If you really don't want to limit jobs' memory use, the resource manager can at least adjust the oom_score.