3 ms·
I don't mind doing the engineering. But penetrating the veil of some of these things, ill-defined as they are, with work on certain areas being sporadic and tak
by mickeyp 4y ago
I don't mind doing the engineering. But penetrating the veil of some of these things, ill-defined as they are, with work on certain areas being sporadic and taking place of years or decades, makes it especially hard.
In fact it's the _absence_ of choice that hurts more than it is a surplus of choice.
As for OOMkiller: I think we're both wise enough to know that experimentation is a large part of what any team that consumes gobs of RAM do. So talk of a "resource manager issue" is all well and good in prod when you should have a reasonable handle on what's used by what and for how long. Less so when you're scaling a model --- say a monte carlo model -- to a larger number of simulations in development and you're testing things. When you're paying an awful lot of money for hardware you start to count (or you should, if you respect your company's money) the costs of things like this.
Regardless, the OOMKiller will indeed reap stuff; and more often than not, it'll pick something shouldn't (it's probabilistic and wrong as much as it is right) and that can cause headaches.
As for NUMA: I'm not blaming the kernel for anything :-)
And hypervisors are sometimes a given, and not a choice. We don't all get to pick our hardware, nor what hardware is made available to us.
- gnufx 4y agoDefinitely there's a lack of shared wisdom about tuning parameters (along with other things in HPC, at least in the circles I see), and too much churn with Linux v. userland (e.g. cgroups). Actually, I was forgetting the stupidity of the memory cgroup invoking the OOMkiller, rather than giving ENOMEM, if you purely use the cgroup for memory accounting -- a sensible-looking change from OpenVZ was rejected. That means there's probably no indication about what's happened unless the job is correlated with the syslog message. As I don't get to do that stuff any more, I haven't looked for a way to hook it now. There's obviously a window in which it can fail, but the resource manager can track the job's PSS, as well as at least ulimiting data and stack. You certainly don't want the job to start paging, if you have swap -- that definitely wastes the resource. If you really don't want to limit jobs' memory use, the resource manager can at least adjust the oom_score.