7 ms·
(The article is from 2018.) I think the article is strongly from the point of view of what to do "in production". If you have a bunch of servers with known spe
by thxg 4y ago
(The article is from 2018.)
I think the article is strongly from the point of view of what to do "in production". If you have a bunch of servers with known specialized workloads, I can believe that enabling swap is good for efficiency. You can run more tasks and get closer to the memory limit while remaining safe. Coming from a Facebook employee, this makes sense.
However, for individual development machines, the pragmatic situation is completely different. The truth is that in practice running out of memory stems from two types of human errors:
1. You wrote some code that leaks lots of memory quickly.
2. You used Google Chrome.
In either situation, still in 2022, Linux reacts in the following way:
A. With swap, your system becomes completely unresponsive for longer than you have patience for, so you power-cycle it.
B. Without swap, the culprit is immediately identified and killed. Your system is perfectly usable, aside from a killed process.
Note that depending on your distro, the good case B. may have become a bit harder to achieve recently: systemd-oomd now tries to eagerly detect OOM situations before the kernel, and when it does, it kills the offending process' whole cgroup! If you have gnome or kde, you're good, but otherwise, this may terminate your session. This can be fixed with "systemctl mask systemd-oomd.service"
- iforgotpassword 4y agoFor me it works well with swap. I'm playing an indie game that suffers from a slight memory leak. After 4 or 5 hours the oom killer comes along and puts an end to it. Everything else gets pushed to swap before the game is killed. No unresponsiveness.
- simiones 4y agoI think case B is not so cut-and-dry. If the code is leaking slowly enough, you will encounter case C, where you don't have swap and the kernel starts swapping out your code pages, and then your computer slows to entirely completely unusable levels for far longer than in case A. Especially if you swap out to a relatively low latency disk (SSD), swap is much better than risking state C.
- deleted 4y ago[deleted]
- charcircuit 4y agoUsually in B my computer will lock up before the OOM killer is ran. I spam alt + sysreq + f and then wait 30 minutes hoping that the OOM killer will eventually kill a process and I pray it doesn't take down X and close everything I had open. I am not running systemd-oomd or any other user space oom killer runner though.
- rcxdude 4y agoB. Without swap, the culprit is immediately identified and killed. Your system is perfectly usable, aside from a killed process. This is not my experience. I ran my desktop for a few years on this theory, and in practice I found the system behaviour far worse when the system ran out of memory, it would lock up completely as you mention with case A. With swap enabled I would eventually reach an unresponsive state, but not immediately: there would be a gradual slowdown in which I could take action to resolve the problem. I believe this is because of the page thrashing of file pages of the executables running (point 3 mentioned in the article). In practice if you want behaviour B you need to run something like earlyoom which kills processes before the kernel starts thrashing the disk.
- kijin 4y agoThere's also B-2. The kernel kills an unrelated process that happened to request a bit of RAM at the moment of OOM. The system becomes responsive for a while, and then the kernel starts looking for a new scapegoat. Which might be the same as the old scapegoat that got automatically respawned by your monitoring tool. This poor program keeps crashing for no reason and you're tearing your hair out trying to find out why. :(
- lamontcg 4y agoYes, in my experience those are the actual tradeoffs. With swap things slow down and that alerts you to the problem and you can go manually OOM kill the right thing. Without swap the kernel "randomly" kills the wrong thing without fail, often leading to a system that you have to reboot to get back into a sane state, and leading to a small industry of trying to tune the OOM killer to never do that. Back in my day (which was a long time ago) we did tune the OOM killer in prod to hit only the right processes first (e.g. apache httpd or whatever software processes were deployed on the server) and that would usually lead to self-recovering behavior where one bad request that caused an OOM in one proc would be killed and the server would recover. That was only in prod though where we understood exactly what software ran on which instances. So, I'd tend to suggest running with little to no swap in prod and tuning the OOM killer because you know what processes are likely to be the issue in prod. While on a desktop/laptop or something its better to have swap because the random workloads you throw at that are going to make tuning the OOM killer impossible. I appreciate the argument that you should still have some swap in prod for paging under utilized anon memory, but I'd like to see some solid numbers about that vs. running swapless and to weight that against the fiddliness of managing the swap files. I suspect in the majority of cases you're going to see that it doesn't make any difference in actual performance numbers, and that you shouldn't be running so close to the edge that it would matter. But measure for your own situation. I was also running in a mostly HDD era not SSD so things may have changed. That probably suggests less of a penalty towards swapping though and the right answer for desktop/laptop loads to be using swap to avoid the randomness of the OOM killer. That may lead though to using more swap in prod since it may degrade and recover more gracefully these days instead of the absolute catastrophe that swapping to the HDD was back in the day.
- planede 4y agoYou are right. I wonder if it would make sense for distros to put processes that are critical for interacting with your computer in a cgroup with 0 swappiness, so your computer remains responsive, even if some processes gets dog slow. Some candidate processes: * ssh-server * X/wayland * Desktop environments * Maybe some terminal emulators and shells? The tricky part is that you probably want at least a shell that remains responsive, but you probably don't want to set their child processes' swappiness to 0. There could be some whitelisted system utilities that are known to not consume a lot of resources, so you can list processes conveniently and kill the appropriate ones. I assume Windows also treats some processes specially so it can remain responsive, and I think now with appropriately set cgroups Linux could have the same capability.
- mort96 4y agoThe only real solution (at least until and unless the kernel OOM killer is tuned to be massively more aggressive, which I doubt will happen) is to run a userspace OOM killer. If you don't like systemd-oomd, there are many alternatives, some which even show a desktop notification when you're dangerously low on memory and when it actually kills processes. Maybe it would be interesting to see if there could be better kernel APIs for things like userspace OOM killers; ideally, we'd want to guarantee that the userspace OOM killer is always prioritized in low-memory situations, and ideally it'd be possible to install low memory event listeners into the kernel rather than to poll.
- AnIdiotOnTheNet 4y agoI disagree. The real solution is to do away with the need for OoM killer in the first place by turning off overcommit (in its current form anyway) and fixing the broken programs that rely on its behavior.
- shawnz 4y agoAnd instead force every application which needs to store large amounts of temporary data to implement its own swapping mechanism? Don't you think we will end up with a lot of even less optimized swapping systems that way?
- silon42 4y agoIf you need swap, you need swap.... You shouldn't borrow it from other programs executable pages (without proper accounting).
- mort96 4y agoIn practice, programs crash when they receive a null pointer from malloc (either through throwing an uncaught exception, or through `if (!ptr) { abort(); }`, or through dereferencing a null pointer). So even if your solution was realistic, it would just entail killing a random process, and it would prioritize killing an essentially random process. When you reach OOM situations and need to kill processes, there are probably better heuristics than "kill whichever process happened to allocate memory after we ran out".
- lazyier 4y agoThere are lots of times I have ran out of memory on modern desktops. Even with ones that are relatively large mounts of RAM, 16/32 GB. The argument for swap is even higher on desktops then Facebook servers because unlike servers you lack the predictable work flow. And you are going to run into more situations were you have memory leaks or processes that are running that you don't use for extended periods of time. Not having swap on the desktop leaves you with a unoptimized system. > B. Without swap, the culprit is immediately identified and killed. Your system is perfectly usable, aside from a killed process. No. This is never been my experience. The normal experience is: OOM killer kicks in, something strange happens, and user is left confused as to why what happened happened. It requires sysadmin skills to go in and interpret dmesg to understand what happened, and even then it is likely to be very unclear without additional information. > Note that depending on your distro, the good case B. may have become a bit harder to achieve recently: It's always been shitty. It hasn't "gotten worse". It's always been bad. OOM is a emergency situation and still leaves your system in an unknown state.
- jjoonathan 4y ago> No. This is never been my experience. The normal experience is: OOM killer kicks in, something strange happens, and user is left confused For linux servers and junior devs, yes, but for my parents on mac desktop, getting a "you are almost out of memory" dialog is simple and actionable while paging hell is mysterious and not actionable.
- gnulinux 4y ago> With swap, your system becomes completely unresponsive for longer than you have patience for, so you power-cycle it. My entire life using linux (>15 years) never had a situation where once swap caused system to slow down, system eventually got back on feet. Usually, swap causes other apps to slow down, swap and put system in a state it can't recover from. Maybe I'm mistaken but my rule of thumb is to never use swap unless I specifically need it. This means, for a default desktop OS, I would never enable swap. If we run out of memory, apps should be killed.
- yabones 4y agoA much better way to handle desktop memory is to have both a generous amount of swap on fast storage, both to allow hibernation AND to allow a bit of wiggle room if things fill up too fast, and also installing and enabling earlyoom(1) with a high but effective threshold to prevent a lockup. Sure, it can cause things to break when using lots of memory, but it never lets you completely stall out. There are now configurations within systemd for this, but I'm simply more familiar with earlyoom daemon. https://manpages.debian.org/bullseye/earlyoom/earlyoom.1.en.html https://manpages.debian.org/bullseye/earlyoom/earlyoom.1.en.... https://wiki.archlinux.org/title/Improving_performance#Improving_system_responsiveness_under_low-memory_conditions https://wiki.archlinux.org/title/Improving_performance#Impro...