6 ms·
PostgreSQL and the OOM killer: Why we use strict memory overcommit
- Bender 3mo agoThey allude to this in the article but I would emphasize caution when using mode 2 especially if one has already adjusted overcommit ratios as one can prevent forks. Test this in a QA/Perf environment first, also testing the restart of all applications. Load test and do full QA tests before deploying to Production and even then when deploying to production I would just dynamically change the setting via app deployment scripts until confidence is high instead of putting it in the sysctl config files. I've gone through this exercise in the past on much older kernels which they cover as well and just me personally I ran into less issues by leaving overcommit to 0 and just dropping the overcommit ratio to 0 and setting the oom_score_adj for programs as high as 1000 if I wanted vmscan to leave them alone and of course using the Redhat formulas for setting vm.min_free_kbytes, vm.admin_reserve_kbytes, vm.user_reserve_kbytes. And of course be vigilant in disallowing app owners from using every last bit of memory.
- Bender 3mo agoCorrecting a rather significant typo: setting the oom_score_adj for programs as high as 1000 should be -1000 to be left alone. 1000 would make it a prime candidate for an OOM kill. Positive integers should be used on sacrificial superfluous programs. [1] As an example OpenSSH sets the sshd to -1000 by default. [1] - https://man7.org/linux/man-pages/man5/proc_pid_oom_score_adj.5.html https://man7.org/linux/man-pages/man5/proc_pid_oom_score_adj...
- szmarczak 3mo agoI have disabled overcommit both on Windows and on Linux. I hate having random programs being killed. Unfortunately, many programs commit 2x memory than they actually use. Often I see ~32GB committed and ~16GB resident.
- nok22kon 3mo agohow exactly did you disabled it on Windows? I dont think it has an option for that.
- szmarczak 3mo agoSettings -> View advanced system settings -> Performance (Settings) -> Advanced -> Virtual memory (Change...) -> No paging file
- 0x1d7 3mo agoThis is almost always a bad idea. If no memory is available where a page file would make a difference, this leads to application crashes instead. A crash is (usually) worse than paging. Certain applications, Photoshop being the historical example, will outright fail to run with no page file present.
- szmarczak 3mo ago> this leads to application crashes instead Same happens if the page file is full. In that case, why don't those programs use disk directly instead? No such problem would've ever occured if programs hadn't allocated more than they actually use.
- 0x1d7 3mo agoYour argument falls flat when a page file can be multi-GB and automatically grow. And if your application admin was competent, memory monitoring would be part of the application monitoring stack. An application that grows in such a way (besides having backing stores for memory-mapped files, as well) will often perform so poorly that it requires addressing (adding RAM, looking for application faults, etc). A page file is insurance, one that can last you much longer than available system memory.
- szmarczak 3mo ago
- leononame 3mo agoThis has bitten me multiple times. The problem I have is that at work we deploy the application (written in Go) and PostgreSQL on the same machine. The backend app allocates a lot of virtual memory, and initially we had overcommit to 0 (heuristic). This caused crashes on big queries in PostgreSQL and we set it to 2. The whole system became a bit unstable because the backend would still allocate a lot of virtual memory and at some point we ran into errors when allocating. For now, we have overcommit_ratio set to a value that is stable from experience, but there really seems to be no silver lining. Go is very happy to allocate a lot of virtual memory, but so are most managed languages. The best solution would probably be to host the backend and the database on separate servers.
- hilariously 3mo agoYes, it would. Basically every serious database tries to allocate everything and more - back in the day we'd just allocate VMs on the machine even with the overhead because knowing it cannot leave its constraints and would work within them was worth the cost.
- guenthert 3mo agoThere are many reasons to use a dedicated host (or VM) for a DB server, but if only the accessible memory needs to be limited a container is the simpler, more efficient tool. Said that, I would expect to be able to configure how much memory a DB process is allowed to allocate. I remember distinctly that PostgreSQL allows such. But of course both can be configured simultaneously, a belts&suspenders approach if you will. Whether failed transactions are actually so much more desirable than a OOM-killed process isn't quite obvious, but it might be easier to troubleshoot.
- xyzzy_plugh 3mo agoI'm not sure if you are aware but there are relatively recent environment variables you can set to help contain Go memory to a fixed size. GOMEMLIMIT works very well if you set it to around 90% of available memory as a rough heuristic. You should definitely profile your application to fine tune this number (e.g. if you link with C libraries that hold large memory pools then Go doesn't account for that) but also to identify sources of spikey/leaky allocations. For example, encoding/json is notorious for it's inner sync.Pool hanging on to outsized buffers. There's usually a lot of low hanging fruit. In my experience Go can be extremely stable in terms of memory footprint at both small (~O(1MiB)) and large (~O(256GiB)) scales, and it takes only a small amount of effort. As far as GC languages go, it is by far the easiest to work with.
- swordlucky666 3mo ago[flagged]
- ozgune 3mo ago(Ozgun from Ubicloud) I agree with the blog post's technical contents, but I feel we came across too strong in the title. For Ubicloud as a managed Postgres provider, we use strict memory overcommit. Our experience with operating Postgres at scale taught us that it's better to enable this than going with the defaults. However, I can see many other scenarios, where using strict memory overcommit would have unanticipated side-effects. That's why Linux doesn't go with strict memory commit as its default.
- furkansahin 3mo ago(Furkan, submitter) Hmm, I haven’t thought about that. I updated the title to better reflect Ubicloud Postgres' position.
- geraldwhen 3mo agoIs this an AI response?
- furkansahin 3mo agoNo.. But I have been using a lot of AI, recently. It might have impacted how I form my phrases? maybe?
- sisve 3mo agoNah, your doing great, you just reflect on you own position and adjust it. People are just (too) suspicious when people are not locked in their ways and are nice instead of hostile
- dang 3mo ago(previous title was "PostgreSQL and the OOM Killer: Why You Must Use Strict Memory Overcommit", if anyone is wondering.) Thanks for updating it here as well!
- otterley 3mo agoI think this is also a good lesson on why it's best to isolate mission-critical services like databases on their own compute nodes.
- deleted 3mo ago[deleted]
- adamors 3mo agoI read this article about 3 weeks ago when this bit me. Really great write-up, some tricky details.
- chiply314 3mo agoNothing worse than memory management on Hyperscaler VMs which do not use Swap :| Took k8s ages to get Swap support. We lost something when we accepted that Hyperscalers just tell you to use more moemory. It was shitty 5 years ago and today especially after the ram price increases
- ValdikSS 3mo agoMy guess would be: it's because memory management before MGLRU was really not good and required different userspace solutions and tinkering. You either get killed with OOM (no swap) or got into thrashing (swap). And now, with PSI + MGLRU, situation is much better, but there are still missing features/subsystems which would be nice to have. For example there's no simple way to lock memory mlockall-style to ensure that rarely used daemon would not face long no-cache-latency upon accessing the first time after long idle time.
- wongarsu 3mo agoFor once, Microsoft's decision to just not do overcommit in Windows seems sensible
- man8alexd 3mo agoWindows doesn't need to fork, and you can't fork a large process without overcommit.
- layer8 3mo agoOne could rephrase the parent that Windows did the sensible thing by not using the fork model.
- fpoling 3mo agoMicrosoft just followed VAX/VMS that does not overcommit. And there is a noise on Linux mail lists to implement process builder pattern which VAX had like 50 years ago…
- baq 3mo agoLinux vm defaults are legit insane in 2026. - system dies under memory pressure (regardless of swapping, actually not having swap makes it worse which should be common knowledge by now) - system dies under disk pressure even if there are tons of free memory (this one is fun to diagnose) - system can technically not die, but render itself useless (or worse) under memory pressure by the oom killer - memory compression of any sort is not enabled - ... Both Windows and macOS do so much better out of the box for essentially any workload.
- koverstreet 3mo agoThe mm people are increasingly hostile to any method of handling OOMs (like, just failing the allocation) besides the OOM killer - it's become very dominated by the hyperscalars and cloud vendors. Working around mm nuttiness is a frequent source of frustration.
- layer8 3mo agomm?
- callahad 3mo agoI believe koverstreet is referring to the Linux kernel Memory Management subsystem (https://docs.kernel.org/mm/index.html https://docs.kernel.org/mm/index.html)
- deleted 3mo ago[deleted]
- cmurf 3mo agoKernel oomkiller is useless for the desktop. It only cares about kernel survival. There's user space oom options, like oomd, earlyoom ... but I'm not sure any heuristic really knows what user space program to clobber. Resource control via cgroupsv2 sounds better for both desktop and server use cases, differing by what processes receive minimum resources. But IO isolation remains elusive. Some suggest a database of hardware capabilities, I wonder if a cheap estimator of throughput and latencies could do this dynamically. Anyway, I'm not convinced user space should care about or rely on the kernel oomkiller. Well before it even considers the situation human interactivity with the system was lost.
- man8alexd 3mo agoMode 0 (Heuristic) is described incorrectly. All this complex heuristic was removed almost a decade ago. Currently, the kernel refuses a single allocation that exceeds the physical memory. That is all. The article ignores the proper modern solution to prevent OOM killing of critical processes - OOM Score Adjust. Tuning CommitLimit manually is an archaic, imprecise, and error-prone way to handle memory limits, only suitable for single-process workloads that can handle ENOMEM properly. It completely ignores dynamic file page cache memory allocation. You still can get OOM if you get unusually high file activity. On the other hand, under low file activity, it wastes memory on the same page cache, because it can't be reclaimed without memory pressure, and memory pressure can't be created because workload hits ENOMEM earlier. Don't use strict overcommit.
- fdr 3mo agoThe key thing is Postgres does handle enomem well and does a nice rollback rather than crashing the server and entering crash recovery. It’s one of the few programs that does. Exceptions for the exception. Even a revised heuristic that only spots large, individual allocations is not going to do the job. Oom score adjust also doesn’t do the job: because the only interesting workload is Postgres, if a backend does a page fault that needs memory, who dies? Another sibling Postgres, almost certainly. Then postmaster does crash recovery, which most would rather avoid. High performance databases with distant checkpoints can take a while to come back up.
- man8alexd 3mo agoThey do have sidecars like prometheus, node_exporter running alongside Postgres and they include them in their MemoryLimit calculations.
- IshKebab 3mo agoThere's so much great stuff here. First, Linux's default memory management strategy is bonkers. OOM killing rarely actually works in my experience, at least on desktop. It takes ages to kick in and usually the system just freezes and you have to hard reboot. I've experienced this on every Linux system I've used, even my current one with 128GB of RAM and 64GB of swap, so don't say "it works for me". Windows and Mac do not have this issue at all, so clearly it's possible to do it better. Has anyone tried using strict overcommit on desktop Linux? Second, this bug is a great counterpoint to those annoying people who naysay Rust with "but not all bugs are memory safety bugs, what about logic bugs? huh?". Rust code would not have had this bug.
- man8alexd 3mo agoYes, many have tried to use strict overcommit on the desktop. It is a good footgun. https://unix.stackexchange.com/a/797888/1027 https://unix.stackexchange.com/a/797888/1027
- IshKebab 3mo agoThat doesn't really explain why it is a footgun.
- man8alexd 3mo agoThe last paragraph: > On the modern desktop, where programmers don't care about failing malloc(), disabling overcommit is shooting yourself in the foot. As you can observe, the memory allocations start failing long before the memory is exhausted.
- 10000truths 3mo agoI'd be interested to see a Linux distribution whose entire shtick is to run well-behaved under a kernel with overcommit disabled. But it would be a huge undertaking. Besides the obvious issue with fork(), there are a lot of programs and libraries out there that implicitly rely on overcommit due to not checking malloc() for failure.
- man8alexd 3mo agoThere is some kind of illusion or myth that strict overcommit solves memory management issues.
- silon42 3mo agoNot before userspace software is fixed, it doesn't.
- man8alexd 3mo agoYou can't fix the fact that you can't predict the future. The software will always allocate more memory than it needs because it can't predict its future resource usage, so limiting memory allocation with strict overcommit is meaningless.
- silon42 3mo agoI disagree... it should not allocate much (like 2x) more than it needs right now... (allocating virtual memory, but not committing is different, and should be handled with MAP_NORESERVE (or similiar)).
- zbentley 3mo agoSigh. Malloc failure should have had to be trapped with a signal or something, not just a return status. I know, I know, threads and nesting handlers make that hard and historical precedent makes it impossible to retrofit, but I can still dream.
- mono442 3mo agoThe problem with disabling the memory overcommit is that then the RAM is wasted. That can be worked around with setting up swap but then the disk space is wasted.
- man8alexd 3mo agoPeople like to reinvent things that they are not aware of. Original BSDs used to use strict swap reservation - every anonymous memory page had to have an associated swap page. You had to have the swap 2x of RAM to allow large processes to fork - otherwise you would get an "out of swap" error. FreeBSD implemented overcommit around 2000, I think version 4.x or 5.x.
- frollogaston 3mo agoWhy are programs even allocating memory that they don't use?
- man8alexd 3mo agoBecause no one can predict the future, and they don't know how many resources they will need.
- frollogaston 3mo agoI mean they can just malloc when they need more, what am I missing? Unless this is about JVMs which might preallocate a ton for their heaps and not use it
- swappie 3mo agoApps could, yes. But it’s about 1000x slower to call the OS to allocate a new page, than to just use a pointer to preallocated space.
- senderista 3mo agoThe proper way to handle OOM is to do what mature databases do: implement your own memory accounting, use only your own allocators integrated with the accounting system, and ensure that every allocation path can recover from OOM. Easier said than done.
- man8alexd 3mo agoMariaDB recently implemented memory PSI monitoring but failed with that in a curious way and disabled it afterwards by default. The failure is that under memory pressure, they flushed the entire InnoDB buffer pool.
- zbentley 3mo agoThe issue is that there’s no generally correct behavior. Should a database under memory pressure stay up at all costs even if it becomes unusably slow (by e.g. nuking 99% of the buffer cache)? Or should it crash/failover hard with a likelihood of potential recovery afterwards, even if it technically could have stayed up? Something in between? There’s no correct-in-general answers to those questions. This is a hard problem due to context dependence; that’s why there are so many knobs.
- man8alexd 3mo agoIn this specific case, the correct behaviour would be to drop a part of the buffer pool until the memory pressure is gone. The context-dependent question is how much and how fast to drop. The current implementation drops to a single configurable level but I suspect it could have implemented better heuristics.
- tomlow 3mo ago[dead]
- ralferoo 3mo agoI have a couple of points, not really sure if they should be in one post or not, but whatever. Firstly, if you take what's written at face value, it seems there's a serious logic error in the OOM killer. If it genuinely counts shared memory against a process, then its logic is wrong because killing that process wouldn't release that much memory, it'd need to kill all the processes sharing that memory to release it. So, maybe it should ignore shared memory in its calculations, or weight them by number of processes sharing it, or whatever. The other issue is kind of true, but shows a stubbornness from the developers to change how they approach the problem. It's true that if any task could be partially through updating shared memory when it is killed, then all bets are off as to the state of that memory. However, if that's the case, then there should already be some kind of locking mechanisms in place to prevent multiple processes updating the same pages anyway. Probably the current solution is: something gets locked when modified; every other process would need to spinlock until it's released; locking process is killed; everything else is stuck; another PG thread notices the child has died and kills everything else; on restart the DB has to be recovered. A different solution could be: give every process its own private part of the shared memory for when it starts a transaction; have an indirection table from page number to shared memory page; for every page that needs to be modified, a new page is allocated from the freed pages list; that allocation is recorded in the private part of the shared memory along with the page number it's replacing; the old page is left untouched and copied to the new page along with any changes; the change list is terminated; then as the last step we update the indirection table for every page that was modified. If the process was killed at any point, we can either roll back the entirety of the transaction (freeing every allocation it made for replacement pages), or if the list was marked as terminated, finish off updating the indirection table with the changes required and marking the original pages as unused. At that point, the lock can be released knowing that shared memory is still entirely consistent. Some of that is probably happening anyway if postgres supports reading from tables concurrently with an active write transaction on the same table, in which case the logic on how to mark those now freed pages as still in use is required. In that case, each process can also maintain a list of pages it has marked as still being used in case a reading process is killed off.
- man8alexd 3mo agoAbout shared memory included in the memory size calculations https://lkml.iu.edu/1902.2/05674.html https://lkml.iu.edu/1902.2/05674.html