5 ms·
That sounds wild. How else can it possibly fail than by failing malloc? Where do the intermittent failures occur, if not in the return value? Could it be some p
by dataflow 2y ago
That sounds wild. How else can it possibly fail than by failing malloc? Where do the intermittent failures occur, if not in the return value? Could it be some part of your code isn't checking the return value for NULL?
- usefulcat 2y agoIIUC, the way overcommit works on Linux is the memory returned by malloc() is not actually mapped to physical memory until it is accessed. So malloc may succeed and then at some point later you may get a segfault or similar when something actually tries to access that memory.
- cyberax 2y agomalloc() will not return unmapped memory for small allocations. Your process might get killed if malloc() fails to get more backing pages.
- dent9876543 2y agoDo you mean that my process might get killed [by the OOM killer] if /another process/ fails to get more backing pages? Or are you suggesting that there is some other failure path that results in my process getting killed if my own malloc() fails to get pages? (That seems wrong. Should simply return NULL in that case, no?)
- cyberax 2y ago> Do you mean that my process might get killed [by the OOM killer] if /another process/ fails to get more backing pages? It can be your process or some other process, depending on the OOM killer weights and heuristics. > Or are you suggesting that there is some other failure path that results in my process getting killed if my own malloc() fails to get pages? (That seems wrong. Should simply return NULL in that case, no?) No, malloc() in Linux will never return NULL for small allocations. What happens is that glibc will try to get more pages for its heap by calling mmap(). The call to mmap() will result in the OOM killer waking up and freeing some memory. It then can kill some other process, and your process mmap() will succeed, or it can kill your process, and in this case mmap() never returns.
- dataflow 2y ago> IIUC, the way overcommit works on Linux Doesn't vm.overcommit_memory == 2 disable overcommit? That's the situation we're talking about...
- usefulcat 2y agoYou're correct; for some reason I didn't interpret the question in that context. I was thinking about when vm.overcommit_memmory is 0 or 1.
- cyberax 2y ago> That sounds wild. How else can it possibly fail than by failing malloc? By killing your process and/or by failing fork()s. In general, malloc() in Linux will never fail for small allocations.
- fulafel 2y agoWith overcommit disabled, fork may fail, but your other claims are backwards: malloc can fail, and processes aren't killed.
- cyberax 2y agoI suggest trying that :) You'd be surprised. glibc allocator, like pretty much any other allocator, maps pages in advance without touching them. It's quite likely that you'll get OOM killed when you first try to touch them, rather than during mmap() for new pages.
- fulafel 2y agoAssuming you've also set overcommit_ratio to a conservative value, this would mean that the kernel overcommit=2 setting is not working, since mmap-ing writable pages counts as allocation from overcommit POV. It's a kernel sysctl, so I wouldn't expect it to change glibc's behaviour. There don't seem to be any open bugs about this in the kernel bugzilla. But do tell if you have found a bug in this area. (This is documented in https://www.kernel.org/doc/html/latest/mm/overcommit-accounting.html https://www.kernel.org/doc/html/latest/mm/overcommit-account... )
- ploxiln 2y ago> malloc() in Linux will never fail for small allocations That's in kernel-space, not user-space processes. If you have overcommit disabled, it will be user-space processes making syscalls mmap() or sbrk() to get more heap memory (in which small allocations reside), that fails with an error return code, I assume. But I guess it could also fail to auto-grow the main thread stack, and presumably get a segfault. If the kernel needs to allocate more memory for a fork(), that syscall can fail. But if the kernel needs to allocate memory for some other internal purpose, probably back to the OOM killer, I guess ... hopefully this would be very rare?
- kevingadd 2y agoFrom what I saw when running into this regularly, it seemed to be code that wasn't checking malloc return values for NULL, or code that didn't have well-defined behavior when out of memory. So I'd get weird error messages from random libraries being consumed by the applications I was using during execution, or individual operations in a larger task would fail but it would keep on chugging along.