31 ms·
The Linux kernel's inability to gracefully handle low memory pressure
- waingake 7y agoI'm so happy someone has made a clear bug report here. Because damn, this is a thing.
- throwaway2048 7y agoYep, even with no swap whatsoever, performance is completely trashed (talking even the mouse lags for 30+ seconds at a time) for like a solid 5+ minutes before the oomkiller triggers, with swap, you might as well just reboot because the system will take perhaps an hour to start responding. Linux is completely useless with ram that is almost full in a way that OSX and windows absolutely are not.
- mort96 7y agoYeah, any time I end up in that situation I just hold the power button and prey to the journaling gods. It's a serious issue, I'm extremely glad it looks like there's actually some progress on fixing it.
- minitech 7y agoEnable SysRq and use SysRq+F. https://news.ycombinator.com/item?id=16561746 https://news.ycombinator.com/item?id=16561746
- blattimwind 7y agoLinux absolutely, positively, requires a decent chunk of free memory because the kernel's algorithms simply do not work and in my 15+ years of Linux experience they NEVER worked. If Linux starts thrashing, and it does so easily, then the box takes what surely is an uncomputable amount of time to recover. One way to make relatively sure that this is always the case is to use a userspace daemon like earlyoom. Though to stay fair all desktop OS'es behave badly when put under memory pressure, it's just that Linux is an order of magnitude or so worse.
- maximente 7y agoare any of the BSDs better in your experience?
- myrandomcomment 7y agoWell my memory for BSD is not fully clear, but when I looked at this back in 2008 BSD handled it better then Linux. The king of sorting this was Solaris. It was rock solid. I would still argue that Solaris is better then Linux as a server, but it does not matter as Linux won. Also, I want my Amgia back :)
- ComputerGuru 7y agoI only ditched the Solaris train for FreeBSD when all hope was gone after Schwartz’s reign came to an end. Solaris got so many things right architecturally that are damn hard to shoehorn in after the fact. Until today, I don’t know if any other *nix has sorted out ABI compatibility across architectures, which allowed running applications cross-compiled for another target to run against the memory-resident kernel without virtualization (only instruction emulation when needed). In practice, it meant that you could run “universal” x86 binaries against both the x86 and x86_64 kernels, even calling into drivers (!!), with guaranteed compatibility. The last I had checked, FreeBSD had made it a goal to unify all numeric values of IOCTLs across architectures, but I don’t believe they are there yet.
- floatboth 7y agoWell, a 32-bit Linux binary on 64-bit FreeBSD can call into the GPU driver and render 3D, at least :) But I don't think anything is guaranteed for sure. 32-bit crap is not exactly a priority, haha. On the other hand, syscalls across architectures of the same bitness are exactly the same, minus the newer architectures just not having some historical abominations like sbrk. (IIRC, syscalls are different between even aarch64 and amd64 on Linux, which is just… why?!?)
- 7y ago
- floatingatoll 7y agoDoes the time until oomkiller change dramatically on spinning metal versus SATA SSD versus NVMe, in a swap-off scenario?
- brendangregg 7y agoI don't think that's a fair comparison: Do you normally run OSX with the swap files completely disabled, or Windows with the pagefile completely disabled? That's what this bug is describing. I'd bet things get pretty nasty on OSX and Windows too, if you tried that. Perhaps the real bug is that Linux distros make it easy to run swapless.
- throwaway2048 7y agoThe behavior is even worse if you leave swap enabled, as I already detailed in my post...
- Aaargh20318 7y agoIsn’t iOS basically a flavor of macOS that runs without swap ?
- earenndil 7y agoYes, but it's also a heavily integrated environment that aggressively quits background programs on memory pressure.
- codedokode 7y agoThat is the only working solution with HTML and JS based applications or apps using GC. Applications should save the state and quit (or be ready to quit) instead of using memory in backround mode.
- Aaargh20318 7y agoBut isn't that exactly what the linked article advocates Linux should also do ? Before quitting background applications it first sends them a request to free memory, in a well-behaved iOS program you use this to clean up your caches and ensure your don't use more RAM than you absolutely need. You should also suspend your state to disk when your app is backgrounded so you can just continue where you left off if your app is killed. Many macOS apps also do this, you can forcefully restart a Mac and after a reboot it'll restore your session to pretty much the exact state you left it in, including any open 'unsaved' files. Linux could implement a similar mechanism to signal apps to clean themselves up and maybe a 'save your state, you're about to get killed' signal.
- jolmg 7y agoYeah. I didn't really think of it as a bug at first, but I'm glad someone called it that. I wish the system would just kill the browser or low priority processes instead of freezing everything in an instant.
- magicalhippo 7y agoWhy does it have to kill the browser? Why can't it tell it "nope, no more memory for you" before it's all gone?
- viraptor 7y agoIt can. System-wide it's one extreme or another. With overcommit enabled, you'll pretty much never get refused. With overcommit disabled, you'll get refusals was soon as you reach the max memory, which means lots of mapped pages wasting space. The middle ground is your own config unfortunately - cgroups can limit available memory, but you'll have to set it up by hand.
- Annatar 7y agoAn operating system should never overallocate memory because one cannot build reliable applications and infrastructure on top of a kernel which is lying to the application.
- angelsl 7y agoThen you disable overcommit. For the general case though, because of the way programs have been written, it is easier to have overcommit on.
- rcxdude 7y agoWindows will never overcommit memory. As a result on my PC unless I dedicate a substantial (over 20%) fraction of my drive to swap (most of which will never, ever be touched) I will 'run out of RAM' far before even half of the physical RAM in my PC has been used. This seems extremely wasteful.
- tbrock 7y agoAgree, but this is only a decent bug report. A better one would just make a C program that mallocs more memory than is available. The "open enough tabs so that it crashes part" is like "banana for scale", it is incredible unspecific. You could probably open HackerNews 50x more than espn.com
- curryst 7y agoWhy? The situation in which it occurs is unambiguous. Allocate more memory than is available and watch it fail less than gracefully. What you're suggesting is merely technical gatekeeping; there is nothing to gain from writing a special purpose OOM-causer, no less specifying that it must be written in C.
- emmelaich 7y agoThere are _numerous_ similar bugs reported in the last 10+ years, in redhat's bugzilla, ubuntu, suse and others. Almost all of them get closed because they get old and no-one is willing to do the necessary work to refine the bug report to dependable reproducibility. And that's not surprising really; it's hard, time-consuming work which may be obsoleted on the next kernel version. That said, some have written memory stressors and have been able to crash or stall a machine but it's still a bit hit and miss.
- mdellavo 7y agofacebook's solution https://facebookincubator.github.io/oomd/ https://facebookincubator.github.io/oomd/
- meruru 7y agohttps://github.com/rfjakob/earlyoom https://github.com/rfjakob/earlyoom has worked well for me in the past.
- deleted 7y ago[deleted]
- htns 7y agoIt's even in debian/ubuntu repos now. I hadn't realized that.
- the8472 7y agoandroid has something similar, the low memory killer daemon (lmkd). both use the recently added pressure stall information (PSI)[0] infrastructure in the kernel to determine when the system is overloaded. [0] https://lwn.net/Articles/775971/ https://lwn.net/Articles/775971/
- senozhatsky 7y agoYeah. OOM-kill handling wants to be a silver bullet, sort of. For instance, Linux kernel provides a number of I/O schedulers or net schedulers, etc. to pick from, but OOM kill is "one size fits all". And it doesn't really look like things are going to change [1][2][3]. [1] https://lore.kernel.org/lkml/alpine.DEB.2.21.1810221406400.120157@chino.kir.corp.google.com/ https://lore.kernel.org/lkml/alpine.DEB.2.21.1810221406400.1... [2] https://lore.kernel.org/lkml/20181024155454.4e63191fbfaa0441f2e62f56@linux-foundation.org/ https://lore.kernel.org/lkml/20181024155454.4e63191fbfaa0441... [3] https://lore.kernel.org/lkml/20181023055655.GM18839@dhcp22.suse.cz/ https://lore.kernel.org/lkml/20181023055655.GM18839@dhcp22.s...
- awalton 7y agoFreeDesktop.org solution (being deployed on several GNOME distributions today): https://gitlab.freedesktop.org/hadess/low-memory-monitor https://gitlab.freedesktop.org/hadess/low-memory-monitor Used in combination with compressed swap (ZRAM), it greatly alleviates this problem on the (at least GNOME-based) Linux Desktop. Still, browser really need to do something about the memory problems they're causing. They're Windows 95-level bad at managing their high memory/leak cases - just leave a browser with more than a few dozen tabs open over night. Especially with a tab that does background fetches (e.g. Facebook or Twitter or something with a lot of timer-driven Ajax queries). I assert that if it weren't for browsers, there'd be no memory problems on modern desktops.
- senozhatsky 7y agolkml.org has been unstable for the past few... umm years, so The Linux Foundation runs its own lkml archive - lore.kernel.org/lkml/ Alternative link, just in case: https://lore.kernel.org/lkml/d9802b6a-949b-b327-c4a6-3dbca485ec20@gmx.com/ https://lore.kernel.org/lkml/d9802b6a-949b-b327-c4a6-3dbca48...
- sigjuice 7y agolkml.org is not official
- zfgnu 7y agoThat html code in lore.kernel.org is weird, I wonder how it's generated.
- mort96 7y agoIt's a somewhat common trick I believe. The idea is this; you want newlines inbetween your tags, but if you have HTML code like `<div>foo</div>\n<div>bar</div>`, you end up with an unwanted text node with a space inbetween the divs which changes how the page looks. By putting the newline inside the tags instead of between them, you don't have any unwanted text nodes.
- blattimwind 7y ago> Your disk LED will be flashing incessantly (I'm not entirely sure why). The VM is basically paging all clean pages in and out constantly as their tasks become runnable. A pretty standard case of thrashing.
- throwaway2048 7y agoThis is with swap disabled.
- blattimwind 7y agoSwap and paging are different concepts. The VM can page clean pages (mmap'd files, like most code your system is running) from disk in and out all day if it wants, in fact, when Linux is thrashing, that's exactly what it's doing.
- AnssiH 7y agoYou can think of all your executables and shared libraries on disk as a kind of read-only swap. If the system is low on memory, a page of program code may be dropped from RAM and re-fetched from disk when it is needed again, i.e. when that section of the code is being executed again. (this effect is not limited to program code, though)
- Iwan-Zotow 7y agoAnything mmaped and read-only (SO included) could be and would be ditched
- ploxiln 7y agoIt should have been clarified, because it is not obvious: The kernel can evict memory-mappings of executable files which are currently running. When they jump to a part of the executable that is no longer in memory, it can page that part back in from the executable file on disk. This is pretty cool. But when memory is very low, the kernel will evict practically all user-space executable mappings from memory, and will be reading back in and evicting back out executable file contents on practically every single context switch. It's just trying so hard to squeeze out some space to make its tasks fit in memory and complete successfully. I think this was the desired behavior of big-iron batch-processing in the 90s? Not sure why it has persisted so long. I'm a big fan of linux and this is my biggest pet-peeve.
- cperciva 7y agoFurther to the comments about the pager hammering the disk to read clean pages (mainly but not exclusively binaries) even if swapping is disabled: In many cases adding swap space will reduce the amount of paging which occurs. Many long-lived processes are completely idle (when was the last time that `getty ttyv6` woke up?) or at a minimum have pages of memory which are never used (e.g. the bottom page of main's stack). Evicting these "theoretically accessible but in practice never accessed" pages for memory frees up more memory for the things which matter.
- throwaway3627 7y agoLinux resource scheduling and prioritization and is pretty awful compared to its popularity. TBH, there are very few OSes that get high-pressure resource scheduling and prioritization right under nearly all normal circumstances. The hackaround for decades on Linux is always adding a tiny swap device, say 64-256 MiB on a fast device in order to 0) detect average high memory pressure with monitoring tools 1) prevent oddities under load without swap (as in OP example).
- ddingus 7y agoSgi IRIX nailed this, FWIW. I would have thought some of the IRIX scheduler made it into Linux by now.
- Annatar 7y agoNo way, any sufficiently advanced technology is indistinguishable from magic and IRIX was so very, very advanced. IRIX hasn't been in development for almost two decades now and it's still more advanced in aspects like guaranteed I/O and software management (inst(1M))... What does that say about it and what does it say about the engineers who worked on it?
- twic 7y agoXFS is still better than any version of that ext dreck!
- sinsterizme 7y agoGlad to see this issue raised! My system hangs for minutes sometimes and is very frustrating compared to Windows and OSX which seem to handle out of memory in a much more user-friendly way. Which seems to be: suspending the offending program and letting the user decide what to do from there. I'm sure there's a reason the Linux kernel doesn't do something similar, but can anyone enlighten me? :)
- deleted 7y ago[deleted]
- adamnemecek 7y agoI'm not sure but the assumption might be that there's generally no user to ask as the computer might be a server.
- IshKebab 7y agoRight but if there is a user to ask then it should ask!!
- noncoml 7y agoEhm, and how does the OS know which is the “offending” process? I think you are confusing the issue raised here with your desktop experience.
- brianpgordon 7y agoCurrently the Linux kernel computes a score for each process based on some heuristics. There's a good introductory article on LWN: https://lwn.net/Articles/317814/ https://lwn.net/Articles/317814/
- lazyguy 7y agoYep and it's about as good as just picking a random process and killing it. It's awesome when you run out of memory and you try to log in only to have it kill sshd.
- deepbreath 7y agoI committed the grave mistake of purchasing a laptop with only 8GB ram and I constantly run out of memory as a result. When it happens, I just repeatedly mash alt+sysrq+f until it kills off some chromium tabs and unfreezes my machine. It essentially behaves like one of those extensions that lets you unload tabs. If needed, you can get the tab back by just reloading the page. The machine slows down to a crawl at 96% usage, and freezes at 97% usage (according to my i3 bar).
- wolfgang42 7y agoA few months ago I upgraded my system to 8GB RAM and I don't see the behavior you describe. It does require a little more care when choosing which programs to use[1], but not significantly so. However, you do need to make sure you have swap enabled to let the kernel efficiently manage its resources. It's not unusual for me to have 1-2 GB of stuff in swap; this doesn't affect performance significantly since it's parts of the system that don't need to run, but if you insisted that they all stay resident then it would put a considerable strain on the system in low memory conditions. [1] The big one for me is that I can't run the Atom text editor, Firefox, and a virtual machine all at the same time.
- kevin_thibedeau 7y agoI run 8GiB without swap. It kills processes off quite nicely when hitting OOM as there is no time wasted with paging out to slow storage.
- the8472 7y ago> slow storage Linux recently gained the ability to swap huge pages to NVMe without having to split them. Combined with THP this can be quite fast.
- eikenberry 7y agoHow recent? Did this happen back in the 4.x series or is it new to the 5.x series?
- alexozer 7y agoA couple weeks ago, one of my physical stick of RAM completely stopped working after yet another Linux out-of-memory-force-poweroff situation. No idea if that could be the proper cause, but I do find it a little funny. I just arrived at this thread after my entire system stalling completely at yet another low memory situation. Let's just say I'm extrememly grateful to discover some of these userspace early OOM solutions in this thread.
- jhallenworld 7y agoThis is a very old problem, I used to see it decades ago when making tape backups. Tar would use move the entire disk through the buffer cache so that eventually everything in it was paged out. The classic solution was to use unbuffered versions of the disk device for backups. What I've always thought is that there should be a working set size limit on a process which includes the buffer cache somehow. The idea is that the process may not use more RAM than this size- if it exceeds it, it must either fail or swap out its own pages, not those from any other process. This would fix the problem for tar- it only needs a tiny amount of memory. I think the situation is very similar with the web-browser example. The browser should not be allowed to force all unrelated data to be paged out.
- throwaway8941 7y ago>working set size limit on a process which includes the buffer cache You can control this with cgroups. Plug a process into a separate cgroup and set the `memory.limit_in_bytes` knob to whatever your heart desires. I use it to limit the qBittorrent's memory usage on my machine. `firejail` is very convenient for doing this. If I don't set a limit (30% RAM in my case), it eats up all the memory with a uselessly large file cache, which does not improve upload speeds at all. https://access.redhat.com/documentation/en-us/red_hat_enterprise_linux/6/html/resource_management_guide/sec-memory https://access.redhat.com/documentation/en-us/red_hat_enterp...
- KingMachiavelli 7y agoDoes the Raspberry Pi suffer from this? All but the latest models have less then 4GB of ram & their storage is often slow SD flash (technically it could be fast but most people have cheap SD cards) so it's fits this scenario perfectly. I guess most users aren't pushing a lot into memory like GUI browsers do.
- rocky1138 7y agoIt has nothing to do with the hardware. This will happen on any system with a default Linux install.
- KingMachiavelli 7y agoI didn't mean to imply it had anything to do with hardware rather I intended to point out that there should be a large base of users (owners of RPis) experiencing the issue since the hardware is exactly what you need to reproduce it (limited RAM & slowish disk speed).
- not2b 7y agoThe program using all the memory here is Chrome (or Firefox). It has more information about what is going on than the kernel does. It should be smarter about memory use when it is trying to consume more memory than is available. Perhaps it could page out background tabs to disk or something similar if memory is low.
- Someone1234 7y agoThose were just examples. This issue can be reproduced using any process.
- userbinator 7y agoI'm not sure how he's getting swap even with swap off, but this seems to be the big disadvantage to having overcommit --- the memory allocator won't ever say NO, so an application can keep allocating memory even if that memory becomes uselessly slow to actually access. Then again, this "allocation will never fail" mentality has also lead to applications being written with such an assumption, and when allocations do fail, they crash. (Arguably, that's better than thrashing the rest of the system.) I don't know if the modern browsers will actually stop letting you open new tabs and just give an "out of memory" error instead of crashing, but that's how most Windows programs are usually written --- without the assumption that allocations can never fail, because on Windows, they can.
- quotemstr 7y agoAll memory is conceptually backed by some file. Under memory pressure, the kernel (and this mechanism is the same on any modern kernel) frees memory by writing pages to their backing storage, then discarding them from RAM. There nothing really special about anonymous memory except that it's backed by swap instead of by a named file on some filesystem. On a system with swap disabled, the backing store is still conceptually swap, but since the swap doesn't exist, pages backed by that imaginary swap can't be evicted from RAM. Pages backed by other backing stores can certainly be evicted from RAM, however, and that's how you "swap" on a swapless system. Note that executable code is almost all mapped so that it can "swap" in this way.
- deleted 7y ago[deleted]
- ars 7y ago> I'm not sure how he's getting swap even with swap off It's not data swap, it's executables. Linux knows it can reread the executable from disk if needed, so it uses those memory pages from other things, and reads them in when needed.
- mktmkr 7y agoThis is why one uses mlockall.
- myrandomcomment 7y agoMy biggest issue is that everything gets swapped eventual in favor of disk cache. I know there are setting, but..it is just wrong.
- quotemstr 7y agoI've never liked the approach Linux kernel and userland take to memory exhaustion. Many people confidently asserts that it never happens. The somewhat better-informed suggest that it's unreasonable to write programs that recover from memory exhaustion because unwinding requires allocation --- a curious belief, because there are many existence proofs of the contrary. Then we get a feedback loop where everyone uses overcommit because everyone believes that programs can't recover from OOM, and people avoid writing OOM recovery code because they believe that everyone is using overcommit and allocation failure is unavoidable. And then they write kernel code and bring this attitude there. Memory is just a resource. If you can recover from disk space exhaustion, you can recover from memory exhaustion. I think the current standard of memory discipline in the free software world is inadequate and disappointing.
- tsimionescu 7y agoIt's not just memory discipline that is pretty bad (and not just in the OSS world). I've recently seen several newer languages refuse to deal elegantly with low-level errors. For example, in a Java server application, if one request encounters some buggy code that tries to read past the end of an array, that request will fail, but all others will succeed - this will give a good chance for the system to be usable, and getting a good bug request, with system-generated diagnostics, for the buggy requests. However, in Go or Rust, the same scenario panics and kills the entire process by default - turning a potentially minor bug in some obscure part of the system into a system-wide crash. OOM is obviously harder to deal with (e.g. if one request is using too much memory, there's no guarantee that it won't be other requests actually seeing the OOM errors first), so if we don't even want to deal with the easy stuff, how can we hope to deal gracefully with the hard ones?
- blub 7y agoCompletely agree - languages like Rust prefer to fail fast and hard. It's certainly an easy to understand solution and is getting the program into a well-known state, but it's also low effort and user-unfriendly. They could have done better, but they would need real exception support for that.
- 7y ago
- makz 7y agoSo... running Linux swapless is a thing? How popular is it?
- sandov 7y agoI used to disable swap because it supposedly reduced the life span of SSDs, but I don't care about that anymore — when it dies, it dies.
- yongjik 7y agoKubernetes, for example, doesn't even support swap. Some bug reports say it won't even run with swap enabled, though I didn't test myself. ¯\_(ツ)_/¯
- jillesvangurp 7y agoI don't know about that but it is very common to specify CPU and memory limits for docker containers. Exceeding those automatically leads to the process being killed. The reasoning is very simple: any form of swapping is completely unacceptable on a production server because it randomly and massively degrades server performance. If you have a cluster of stuff and one node is misbehaving like that, you kill it because that is completely unacceptable. If that is a regular thing, your servers are obviously mis-configured in some way and you fix it by provisioning more hardware or tweaking the limits. 12 years ago, before I got a mac, I had a windows XP laptop with enough memory (8GB) to disable the swap file (which world+dog will insist is a very stupid thing to do). This was great and vastly extended the useful life of my laptop. Alt+tabs were instant and I could run e.g. JVM applications with sane heap settings as well as a browser, office stuff and a few small things I needed with zero issues. On the rare occasion that something did run out of memory, it died or I killed it. Laptop disks were stupendously slow at the time; any form of swapping on a slow laptop disk is extremely disruptive. SSDs are much better but there too it tends to be mostly redundant. IMHO most forms of swapping are highly undesirable on both servers and end user hardware. Swapping to free up memory for file caching is simply unacceptable when you can instead just evict cache pages. If you don't have enough memory left to cache effectively, that just means things like memory mapped files will get a lot slower. If something allocates more memory than you have just kill it.
- Animats 7y agoAh, yes, that bug. Few programs can handle a fail return from "malloc", and Linux perhaps tries too hard to avoid forcing one. Most programs just aren't very good at getting a "no" to "give me more memory" Browsers should be better at this, since they started using vast amounts of memory for each tab. I used to hit a worse bug on servers. If you did lots of MySQL activity, so that many blocks of open files were in memory, and then started creating processes, you'd often hit a situation where the Linux kernel needed a page of memory but couldn't evict a file block due to some lock being set. Crash. That was years ago; I hope it's been fixed by now.
- ailideex 7y agoYou provide limited information but it is not clear the scenario you explain is a bug. If too much memory is locked into resident memory with mlock then this sounds like the expected and correct behavior.
- hobbes78 7y agoThen I prefer the unexpected and incorrect behaviour of Windows, which freezes the offending application and continues to be responsive, allowing me to kill it if I wish to do so...
- HugThem 7y agoI witnessed MySQL bringing linux servers down two. In my case it happens like this: I have a long running PHP process that constantly fires away mostly SELECT but also a bunch of INSERT and UPDATE statements and also some DELETEs. Since the DB and the key files do not fit into memory, its all disk bound work. All tables are MyISAM. Like clockwork, this stalls the virtual machine once per day. All I can do is to hard power down the VM and restart it. Afterwards the table data is corrupted beyond repair. Not sure it is related to memory though. Because the memory usage of PHP and MySQL seem to be constant. Most RAM seems to be used by Linux for caches.
- z3t4 7y agoTry doing some rate limiting in order to not cause the dead lock. Should probably also disable write cache. And if it still doesn't work switch to a bare metal machine. And give it a lot of swap and up the swappiness. Swapping is a much better alternative then crashing. VPS providers doesn't like swap because it will tear their SSD disks, so the swap and swappiness is probably preset too low.
- ajyotirmay 7y agoYes, it has been bugging me a lot. My SWAP space remains empty, and my RAM runs out of space. And that situation is frustrating because I can't close my applications or even restart display server to clean up memory. Something needs to be done here for real, otherwise Linux is a nice software
- bArray 7y agoSimilarly, there are many annoying Linux bugs: `pthread_create` can sometimes return back a garbage thread value or crash your program entirely without any way to catch it or detect it [1]. High speed threadding is hard enough as it is, without the kernel acting non-deterministically. Un-killable processes after copy failure (D or S state) [2]. If the kernel is completely unable to recover from this failure, is it really best to make the process hang forever, where your only available option is to restart the machine? I ran into this with a copy onto a network drive with a spotty connection, that actual file itself really didn't matter - but there was no way to tell the kernel this. Out Of Memory (OOM) "randomly" kills off processes without warning [3]. There doesn't appear to be a way to mark something as low-priority or high-priority and if you have a few things running, it's just "random" what you end up losing. From a software writing stand-point this is frustrating to say the least and makes recovery very difficult - who restarts who and how do you tell why the other process is down? [1] https://linux.die.net/man/3/pthread_create https://linux.die.net/man/3/pthread_create [2] https://superuser.com/questions/539920/cant-kill-a-sleeping-process/541493#541493 https://superuser.com/questions/539920/cant-kill-a-sleeping-... [3] https://serverfault.com/questions/84766/how-to-know-the-cause-of-a-oom-error-on-linux https://serverfault.com/questions/84766/how-to-know-the-caus...
- throwaway2048 7y agoThere is now a way to mark processes as OOM-killer exempt https://backdrift.org/oom-killer-how-to-create-oom-exclusions-in-linux https://backdrift.org/oom-killer-how-to-create-oom-exclusion... Part of the issue with processes stuck in D state (waiting for the kernel to do something) is that it is deeply tied into kernel assumptions about things like NFS, NFS is stateless, and theoretically severs can appear and disappear at will, and operations will keep working when it comes back. You can make NFS a hell of a lot less annoying in this regard by mounting it with soft or intr flags, however if the network disappears or hiccups, you WILL lose data (the network is NEVER reliable, in fact the entire model of NFS is arguably wrong to begin with)
- marmaduke 7y ago> the entire model of NFS is arguably wrong to begin with On local networks (everything attached to a single switch) with good hardware, it is reliable, and soft/intr is the worse choice among others. To wit NFS is one of the commonly supported VM storage options (libvirt, VMware, etc).
- garbre 7y agoI've tried OOM killers and thrash-protect. I've tried numerous tweaks to the vm and swap setting. Nothing works. Memory use gets into the 90%s and the system freezes, hard. Nonetheless, I'm surprised someone is calling this a bug. Let's face it, Linux is just not a desktop operating system. It's a server operating system, and it expects that it will be professionally administered and tightly controlled to prevent OOM situations. That OOM situations occur on servers too is beside the point. There are reasons for the linux memory system to work as it does, reasons Linus will yell at you about if you complain.
- hyperion2010 7y agoI use a swap file these days because in the 4 years since I purchased my computers I went from never hitting 32gigs of memory used at the same time, to hitting it once a week. The worst offenders are browsers and the JVM. The swap file saves me from those 20 seconds of distraction when running a variable memory workload that suddenly jumps over the limit and hardlocks the computer for hours on end. If I was doing something important I will wait for OOM killer to maybe reap the evil children, but otherwise I just power cycle the system and add a note to put the swap file in fstab.
- codedokode 7y agoIt is interesting how Android that typically has less memory than desktop systems solves such problems. It kills inactive applications and background browser pages. The program that can save its state is more complicated, but it works better with limited amount of memory. Today there are many applications written using languages like HTML or JS, or garbage-collected languages and unless you can unload them from memory, there will never be enough of it.
- aitchnyu 7y agoWish desktop browsers do this by default, except for pinned tabs or last 20 pages. But then frontend guys make heavy apps with with animated meme loading screens.
- WalterBright 7y agoI've noticed a similar problem with low free disk space with every OS I've tried it on. All kinds of erratic behavior, hangs, etc.
- deleted 7y ago[deleted]
- alexghr 7y agoI hit this bug yesterday on my laptop (16GB of RAM / 1GB of swap) with 2 instances of Firefox (about 60 tabs), Slack, Insomnia (Electron-based Postman clone) and a couple of `node` processes watching and transpiling.. stuff. `kswapd0` was running at 100% CPU, I guess trying to free up some RAM by moving things to swap (the swap partition was full by this point). Luckily I managed to recover the system by switching to another tty and killing kswapd0 and the node instances. Sometimes instructing the kernel to clear its caches helps: `echo 1 | sudo tee /proc/sys/vm/drop_caches` [1] [1]: https://serverfault.com/questions/696156/kswapd-often-uses-100-cpu-when-swap-is-in-use/696185#696185 https://serverfault.com/questions/696156/kswapd-often-uses-1...
- IshKebab 7y agoI wouldn't enable swap on a desktop Linux system. When you run out of memory if you have swap the system grinds to a halt and you pretty much can't do anything to save it, or at least it is a battle. Without swap it just kills processed until there is enough memory, which is what you would have done anyway! I think the main annoyance with Linux here is that in Windows you get to choose what to kill, whereas in Linux it can't really communicate with you (because the kernel doesn't know about such modern things as GUIs) so it had to pick more or less randomly.
- caf 7y agoIt's not really that random. To see what the OOM killer heuristic currently considers its top 5 targets on your machine: for P in /proc/[0-9]*; do echo $(cat $P/oom_score) $(cat $P/comm); done | sort -n | tail -5 For me right now that shows 4 "Web Content" processes (firefox tabs) and a firefox-esr. That seems to check out.
- michaelmrose 7y agoHaving some swap with low swappiness allows the system to page out pages that are unlikely to be used or haven't been used in a long time but aren't backed by a file. In normal conditions this frees up memory for more useful data and helps you avoid getting to perverse conditions.
- jangid 7y agoAt one point we used to brag about using less memory. Common discussions were - "see, my memory usage — while hitting free -m on cli", "see, my kernel size is just 500kb, I chose just the right module" and so on…
- GuB-42 7y agoI wonder how they did with Android? Especially in the early days, not with today's 8GB+ monstrosities. My first Android device was a Nexus One. 512MB of RAM for what is essentially a full Linux system. Able to run a browser and multiple Java apps, all isolated and running their own VM. Task managers often reported near 100% RAM use and things still worked fine. And my understanding is that optimized things further, but given how overpowered phones are today and how bloated apps are, it is hard to tell.
- sshb 7y agoAndroid is described further in this thread https://lkml.org/lkml/2019/8/5/1121 https://lkml.org/lkml/2019/8/5/1121
- known 7y agoI read somewhere Linus Torvalds recommending https://en.wikipedia.org/wiki/Paging#Swappiness https://en.wikipedia.org/wiki/Paging#Swappiness to 90 and let Kernel decide what/when to swap;
- regularfry 7y agoThis is an obvious idea, so I presume there's a reason why it wouldn't work, but what would happen if you had different rules for uid=0 pids and for everything else? If processes running as root were never eligible for oom-killing, and could force mallocs by triggering an oom-kill of user processes as necessary, wouldn't you always be able to recover a thrashing system from a root console? Or is it too hard to isolate console IO from the rest of the system in that situation?
- mschuster91 7y agoMany systems (embedded and phones) don't have a proper root console, and people still run daemons which can leak memory as root.
- regularfry 7y agoYeah, not having a root console on those systems means they're in no better (or worse) a situation than today. I think misbehaving root processes are a different case, though: any misbehaviour as root can have dire consequences, which is why running as root is a Bad Plan in general. Would adding "...and oom behaviour can wedge the box in new and interesting ways" to the list make things worse than they are today?
- malux85 7y agoRead the source code of the OOM killer, the way that it computes the badness score favours userland processes, but root processes must also be considered because it’s possible they misbehave too
- regularfry 7y agoPossible, yes, but the smaller number of root processes and their general higher-risk status implies that a different policy entirely might be better than the same policy with different parameters. Besides which, I'm looking at https://github.com/torvalds/linux/blob/master/mm/oom_kill.c https://github.com/torvalds/linux/blob/master/mm/oom_kill.c and I genuinely can't see where root processes get privileged. I can see a reference to a `root bonus` in a comment, but other than that... maybe my kernel source reading skills are just too rusty.
- nielsole 7y agoAlso, when you run low on memory, iptables can crash on you: https://bugzilla.kernel.org/show_bug.cgi?id=200651 https://bugzilla.kernel.org/show_bug.cgi?id=200651 Poof, no more networking
- ungamed 7y agoNo, just not 'that operation' afaics.
- hagreet 7y agoAfter reading through the comments I would like to know the following: Why are there no priorities? I can't figure it out from the answers. I think with root privileges it should just be possible to say "GUI has higher priority" etc... Then when there is a memory issue you kill some low-priority processes to get the memory back. But whatever any sane person considers part of the operating system because it is the bare-minimum of what is required to do stuff (filesystem, gui, ...) needs to have priority and always be fast. This can be defined by the distribution using the startup privileges. So, why is this so difficult?
- draugadrotten 7y agoRead the comment by idoubtit in the thread below to learn how to prioritize: idoubtit 2 hours ago | unvote [-] Point 3 is wrong. OOM killing is not random. Each process is given a score according to its memory usage, and the highest score is chosen by the kernel. The way to mark priority in killing is to adjust this score through /proc. All of this is documented in `man 5 proc` from `/proc/[pid]/oom_adj` to `/proc/[pid]/oom_score_adj`. http://man7.org/linux/man-pages/man5/proc.5.html http://man7.org/linux/man-pages/man5/proc.5.html
- dingo_bat 7y agoSo Linux cannot handle low memory situation if you disable the low memory management feature (swap). Who knew?
- cyborgx7 7y agoI've been living in denial about this being a linux specific issue until I saw this post. Eventhough I encounter this problem frequently on linux and it has almost never happened to me on OSX or Windows, I've just been telling myself that it was because of the hardware I was using in each case. If they found a way to fix this, where just the browser froze and not the entire OS, it would be a huge improvement.
- Tonitrus 7y agoThe elephants in the room are Chromium and Firefox. They turned the job of displaying an HTML page into a cpu and memory inefficient nightmare. The display of 4KB of information, that is the informational weight of a typical webpage, must not take up 400.000MB of memory and millions of cpu-instructions. Look out for simpler HTML plumbers e.g. Dillo and w3m and help refine their table rendering.
- zwaps 7y agoThis exact bug has been a huge issue for me when I am developing with Matlab. Those are large simulations. Things get swapped around and memory is often close to the limit. Linux then becomes unresponsive, and basically stalls. Theoretically it recovers, but that process is so slow that the next stall is already happening. It is therefore impossible to run large scale Matlab simulations on my Linux machine, while it is no issue in Windows. As far as I can see, Linux is only usable with enough RAM so that it is guaranteed you never run out. I don't know why this has never been an issue, I guess because it is a Server OS and RAM is planable, or very infrequent?
- ri0t 7y agoAdd swap (Windows does that, too) and never use more memory than you have RAM (edit: ..in one process). Storing swap on a quick storage adds to the fun and the price. The one trick, SSDs don't like.
- zwaps 7y agoIsn't swap the standard configuration? I did not set up this server, but I don't see why it would not have swap. Nevertheless, this problem occurs. I'll check. I know the solution is to never use more memory than I have RAM, but that's just what happens - and I know there is no way to "solve" the issue and make it magically run. The issue is HOW it is handled... It's weird that other OSs can deal with this, while Linux needs to be restarted. I think the issue is that only a handful of (Matlab) processes eat up all the RAM, so this "OOM" can not really do anything - there's no use killing off other processes. What should probably happen is to kill Matlab or one of its processes, even though it is in use. I'd be fine with that. Give some out of memory error and kill or suspend the process. At least then we know. Instead, the system just locks up completely (because the other threads keep trying to push stuff into memory), but is not actually dead, so we don't even catch the issue. Also, because EVERY process is essentially stalled, you can't even kill Matlab yourself. Or suspend it and dump the data, which would be useful. No, you have to hard reset the machine.
- Dayshine 7y agoSwap doesn't resolve this on a HDD though. The UI/terminal still locks up, and you still can't recover once you hit the point of thrashing. What really confuses me is that this kernel was developed when SSDs didn't exist, so how on earth did "The system becomes irrecoverably unresponsive if a single application uses too much RAM" get missed?
- 11235813213455 7y agoI had exactly this issue 3 years ago, when I was still using 4G RAM and working with heavy frontend stacks (gulp, webpack1, ..)
- liopleurodon 7y agoI used to have a cgroup just for chrome to limit how much total ram it could use because of this exact thing
- avodonosov 7y agoI increased swap to avoid that on my laptop.
- w-m 7y agoWhen working on Ubuntu 16.04 LTS, this is such a productivity killer. Quite annoyed at the time lost from this behavior, after coming from a Mac. In shells where I run a program that may load a larger data set (e.g. before ipython), I now regularly run `ulimit -v 50000000` to limit the shell's virtual memory to ~50 GB of the available 64 GB on this machine. If the program tries to use more RAM it'll then just die, and not drag down the whole system with it. Works fine, but I really shouldn't have to do this.
- rwallace 7y agoHow are all the people talking about Windows here getting it to behave better? In my experience, when you run out of memory on Windows, the whole machine locks up hard for ten or fifteen minutes while it thrashes the disk before finally killing the offending process. (Admittedly that's on spinning metal; SSD would probably do better.)
- sp332 7y agoHave you tried this with swap space disabled? In my experience it hits a wall very suddenly.
- rwallace 7y agoNo, the above is with the default swap settings. I have tried Windows with swap disabled, but on an 8GB machine, it for whatever reason ran out of memory with only about 4GB of stuff loaded, so I decided it was better to put up with the default settings.
- davidparks21 7y agoOn Windows I could consistently open a 32GB matrix in Matlab with 16GB of RAM on my laptop and perform operations on the matrix. The disk would spin, and it would take 20 minutes to do a simple operation because of the swapping, but I could open it, perform the operation, save, and exit successfully. I could easily background Matlab and do email or other common tasks such as browsing with very little impact to those applications. On Linux Mint that same task locks the mouse and brings the system to its knees, I can't even kill Matlab and would typically resort to a hard reboot. I learned quickly that I can't do the same things on Linux Mint that I used to do pretty easily on Windows.
- rwallace 7y agoFor me, Windows (7, 64 bit) behaves exactly as you report for Linux. I would love to be able to get Windows to behave like it does for you. What version were you using? Did you tweak any settings?
- linsomniac 7y agoThe most annoying thing about OOM is when a process goes crazy and starts using a lot of memory, the OOM killer looks at the system and sees that process is really active, so it kills mysql/ssh/apache/postgres to make room for the run away. I've set up monitoring that pages me when "dmesg" includes "OOM".
- punnerud 7y agoApple have solved some of it by limiting the maximum number of processes before the user get a warning and have to close some of the old. Not sure if they handle memory any better, from my experience; no.
- mort96 7y agoIf there's one process going crazy and using tonnes of memory, what does a maximum number of processes help? We're not talking about a forkbomb here, just one rogue process.
- isodude 7y agoYou can actually adjust the oom_score on those long lived important processes to hinder OOM to kill them.
- linsomniac 7y agoAnd I've always meant to go in and set SSH to be immune to OOM, and deprioritizing others like databases, but I've just never gotten around to it. Looks like oomd could be useful, once we get to kernels that support it in production (looks like 4.20, Ubuntu 18.04 has 4.15).
- isodude 7y agoAnother interesting aspect is to set SSH to realtime. When you log in you set it non-realtime if you work is not important to avoid sinking the server into oblivion by a simple command. It is possible to do when you have shared servers as well, by letting root log in to another port instead and setting that process to realtime instead. Thoughts in my head but never got around to do it.
- C4stor 7y agoI'm seriously impressed by the quality of the discussion generated from this on the mailing list, a great example of online collaboration imho !
- parentheses 7y agoWhat is the solution here? Are you recommending a swap space be created automatically on behalf of the user? One could also use compressed memory (`zram`).
- zzzcpan 7y agoswap on zram (RAMSIZE - 1G) and earlyoom
- igneo676 7y agoWhat are these settings he's referring to? Genuinely asking - I used to run into this issue _all the time_ and even if it's not a default I'd love to toggle some flags and get a responsive system even under low-memory situations
- sunseb 7y agoI fixed this bug myself... I bought extra RAM. :)
- davidparks21 7y agoIn Windows on a 16GB RAM laptop, I've often fired up Matlab, opened a 32GB matrix, and performed a few simple operations on it. In Windows Matlab dutifully chugs away on the problem, the disk spins like mad, and I put Matlab in the background and do email for 20 minutes. This identical use case completely cripples my Linux Mint OS, the mouse hangs, nothing functions, and I've never gotten it to even complete the operation. I just can't operate on a 32GB matrix with 16GB of RAM in Linux, but I can in Windows with relative ease. To me, this is the Linux kernel's biggest weakness against Windows. Most other gripes about Linux (poor power management, poor driver support, etc) belong outside the kernels domain, but this one is a glaring win for Windows over Linux.