3 ms·
> every process is supervised by some other process which will restart it on failure I'm curious if this works in practice for you. The current OOM algorithm
by HerrMonnezza 13y ago
> every process is supervised by some other process which will restart it on failure
I'm curious if this works in practice for you. The current OOM algorithm in Linux sums up the memory usage of a process and all its children. So there are good chances that the restarter process is killed first, and then the main software is killed too (when OOM killer realizes the last kill didn't free enough memory).
This is exactly the problem we're facing at work here: on a computational cluster, users sometimes start wild code that consumes all the memory, but the OOM decides to kill the batch-queue daemon first, because it's the root of all misbehaving processes. We have to explcitly set `oom_adj` on the important daemons to prevent the machines from becoming unresponsive because of a bad OOM decision.
- derefr 13y agoMy "restarter process" is upstart. It's convenient, since the OOM-killer tries to not kill init (for bad things happen when you kill init), so it's a somewhat-safe place to put supervisory logic. One of the better calls Canonical has made, I think. :) Still, in your use-case, I'd definitely recommend only letting users run their "wild code" inside a memory cgroup+process namespace (e.g. an LXC container.) Crash-only systems only work when a faulty component crashes itself before it crashes you. Processes modellable as mutually-untrustworthy agents should always have a failure boundary drawn between them. (User A shouldn't be able to bring down the cluster-agent; but they shouldn't be able to snipe user B's job by OOMing their job on the same cluster node, either.) And on a Unix box, the only true failure boundaries are jails/zones/containers; nothing else really stops a user from using up any number of not-oft-considered resources (file descriptors, PIDs, etc.)
- cbhl 13y agoDo you have any good resources on where to get started going about setting up failure boundaries/jails/zones/containers like this properly? I think it's surprisingly easy to get yourself in the situation where this is a concern for you[0] but you don't know how to solve it. [0] Just run "adduser" and have SSH running, or just create an upstart job, or write a custom daemon that accepts and executes jobs from not-quite-trustworthy-undergrads, or...
- timClicks 13y agoIf you are running Ubuntu, docker.io makes life pretty easy for you to create and maintain LXC containers.
- adobriyan 13y ago> The current OOM algorithm in Linux sums up the memory usage of a process and all its children. It used to sum until Linux 2.6.36.