4 ms·
You're right, I probably should have said "some operating systems". I was generalizing based on my own education, which used Linux. I believe Windows does not d
by offlinemark 6y ago
You're right, I probably should have said "some operating systems". I was generalizing based on my own education, which used Linux. I believe Windows does not do this either, but I'm not sure how many OS 101 classes teach Windows kernel :)
> The advantage is that you find out if you don't have enough memory when you make the system call to request it, not by having the program killed.
I wonder how much of an advantage this is. Memory availability is constantly in flux, and Linux at least, goes to great lengths to reclaim memory before invoking the OOM killer, which is a last resort. What if memory frees up immediately after the application gets the error return, but before the app touches it?
I'd trust the kernel more to manage this than applications themselves, because the kernel can make much better decisions on how to juggle things (evicting caches and whatnot). Plus, we all know how well we all test our error handling paths :)
> (It may be time for paging out to disk to go away. It's a huge performance hit. Mobile doesn't do it. And RAM is so cheap.)
In case anyone is interested in learning more, this is an excellent article discussing the nuances of paging to disk: https://chrisdown.name/2018/01/02/in-defence-of-swap.html https://chrisdown.name/2018/01/02/in-defence-of-swap.html
Re: mobile, does Android disable paging file mappings to disk, which happens even if swap is disabled? (Swap only affects anonymous mappings).
- danieldc 6y ago> I believe Windows does not do this Windows can do both: https://docs.microsoft.com/en-us/windows/win32/memory/reserving-and-committing-memory https://docs.microsoft.com/en-us/windows/win32/memory/reserv...
- wahern 6y ago> I wonder how much of an advantage this is. Memory availability is constantly in flux, and Linux at least, goes to great lengths to reclaim memory before invoking the OOM killer, which is a last resort. Whether or not it goes to "great lengths" before invoking the OOM killer is irrelevant. 1) On heavily loaded systems the OOM killer can trigger regularly, 'causing all manner of havoc because it can non-deterministically kill unrelated processes[1], and 2) a good kernel should go to great lengths to service a memory request, but that merely begs the question if great lengths should involve shooting down unrelated processes. > What if memory frees up immediately after the application gets the error return, but before the app touches it? What if what? What if I try to create a TCP connection to a host which rejected the connection, or the request timed out, but if the request had been delayed just a little longer it would have succeeded? What does that have to do with whether the kernel should start killing unrelated TCP streams? > I'd trust the kernel more to manage this than applications themselves, because the kernel can make much better decisions on how to juggle things (evicting caches and whatnot). Plus, we all know how well we all test our error handling paths :) The kernel has very little information with which to determine which process to kill as an appropriate response to resource exhaustion, presuming killing any process is even the correct response. It's a difficult problem, but OOM only provides one solution that is acceptable to those who wouldn't otherwise even care at the expense of incapacitating the ability of smart software which cares strongly to implement correct and deterministic solutions. Moreover, we're talking about the kernel here. Generally speaking, kernels should provide mechanism, not policy. By heavily and deeply conflating mechanism and policy the OOM killer is a fundamentally and fatally flawed approach for a general purpose kernel. There is no way to "fix" an OOM killer approach without effectively erasing the line between application software and the kernel. That might be fine for embedded systems, but for a general purpose, multi-user system (which in the age of k8s is making a comeback), it's just plain wrong. [1] For most of 2019 a QoI (not correctness) regression in Linux' memory reclamation and page buffer code resulted in JVM processes on our k8s clusters doing heavy disk I/O and generating alot of buffer cache churn indirectly triggering OOM on a daily basis, causing unrelated, twice-removed processes to be terminated as various allocation requests (often internal to the kernel) timed out. Worse, processes could hang indefinitely on various mutexes in the kernel that are taken as part of the "great lengths" Linux goes through. Sometimes very important services would get killed, like systemd or docker. Facebook and Google understand the potential for this problem well, which is why on all their clusters they run their own userspace daemons which try to predict imminent invocation of the OOM killer and throttle or shoot down processes according to their much more sophisticated heuristics and policies. The whole charade of infinite regress is plainly ridiculous, IMO. Almost everybody would be better off with a simple rule that the process requesting an allocation was killed. That's not as good as strict accounting giving developers and processes the option to continue or die (much software does in fact exit on malloc failure, while a not insignificant amount of software could indeed successfully and correctly continue) at the point of commitment, but still preferable. Then a ton of overwrought and buggy code in the kernel could disappear overnight, and smart applications would have a better path forward for achieving more correct, more deterministic behavior. But what we have instead is a horrible hack deeply embedded in the kernel so that mostly mythical programs (i.e. software that preallocated most of system memory but then forked) could run (non-deterministically, of course) on early Linux systems.
- mlvljr 6y agoWould the situation be better in the BSD land?
- offlinemark 6y agoThanks for sharing in this detail. I don't have production SRE experience, so this is fascinating to read. I understand that the OOM killer has caused you quite a lot of pain! I shouldn't have even mentioned the OOM killer. My point was merely: Due to those great lengths, memory availability is highly dynamic. Committing memory up front prevents optimizations where the kernel can reclaim memory "just in time" to allow a mapping to succeed and not require killing anything.