6 ms·
It's all very benchmark-chasing theoretical. In practice performance is more complicated and mmap is this weird corner-case over engineered inconsistent optimiz
by xoo1 6y ago
It's all very benchmark-chasing theoretical. In practice performance is more complicated and mmap is this weird corner-case over engineered inconsistent optimization thing that often wastes or even "leaks" memory which can be used for actually important for performance caches, it's also awful at error handling and so on. I had to literally patch LevelDB to disable mmap on amd64 once, which eliminated OOMs on those servers, allowed me to run way more LevelDB instances and overall improved performance so significantly, that I had to write this comment.
- CalChris 6y agoNeither mmap() nor read()/write() leak memory.
- jstimpfle 6y agoBut they might "leak" it. What parent meant is that as an mmap() user you have no control how much of a mapping takes actual memory while you're visiting random memory-mapped file locations. Is that documented somewhere?
- rowanG077 6y agoThat's called a space leak not a memory leak.
- AnotherGoodName 6y agoThe paging system will only page in what's being used right now though and paging out has zero cost. Old data will naturally be paged out. To put it directly the answer is each mmap file will need 1 page of physical memory (the area currently being read/written). There may be old pages left around since there's no reason for the OS to page anything out unless some other application asked for the memory. But if they do mmap will go to 1 page just fine and there's zero cost to paging out. I feel mmap gets a bad reputation when people look at memory usage tools that look at total virtual memory allocated. I can mmap a 100GB of files, use 0 physical memory and a lot of memory usage tools will report 100GB of memory usage of a certain type (virtual memory allocated). You then get articles about application X using GB of memory. Anyone trying to correct this is ignored. Google Chrome is somewhat unfairly hit by this. All those articles along the lines of "Why is Google using 4GB with no tabs after i viewed some large PDFs". The answer is that it reserved 4GB of 'addresses' that it has mapped to files. If another application wants to use that memory there's zero cost to paging out those files from memory. The OS is designed to do this and it's what mmap is for.
- labawi 6y ago> paging out has zero cost Paging out, as in removing a mapping, can be surprisingly costly, because you need to invalidate any cached TLB entries, possibly even in other CPUs. > each mmap file will need 1 page of physical memory Technically, a lower limit would be about 2 or so usable pages, because you can't use more than that simultaneously. However unmaps are expensive, so the system won't be too eager to page out. Also, for pages to be accessible, they need to be specified in the page table (actually tree, of virtual -> physical mappings). A random address may require about 1-3 pages for page table aside from the 1 page of actual data (but won't need more page tables for the next MB). > application X using GB of memory I think there is a difference between reserved, allocated and file-backed mmapped memory. Address space, file-backed mmapped memory is easily paged-out, not sure what different types of reserved addresses/memory are, but chrome probably doesn't have lots of mmapped memory that can be paged out. If it's modified, then it must be swapped, otherwise it's just reserved and possibly mapped, but never used memory.
- AnotherGoodName 6y agoI'd argue the costs with paging out are already accounted for by the other process paging in though. The other process that paged in and led to the need to page out had already led to the need to change the page table and flush cache.
- labawi 6y agoPaging in free memory (adding a mapping) is cheap (no need to flush). Removing a mapping is expensive (need to flush). Also, processes have their own (mostly) independent page tables. I don't think it would be reasonable accounting, when paging-in is cheap, but only if there is no need to page out (available free memory). Especially when trying to argue that paging out is zero-cost.
- owl57 6y agoAre page tables garbage-collected in Linux? That seems like a potential non-trivial leak source: up to 200MB for your hypothetical 100GB file.
- vlovich123 6y agomadvise gives you pretty good control over the paging, no? Generally I think you can MADVISE_NOTNEEDED to page out content if you need to do it more aggressively, no? The benefit is that the kernel understands this enough that it can evict those page buffers, things it can’t do when those buffers live in user-space.
- jandrewrogers 6y agoNo, madvise() does not give good control over paging behavior. As the syscall indicates, it is merely a suggestion. The kernel is free to ignore it and frequently does. This non-determinism makes it nearly useless for many types of page scheduling optimizations. For some workloads, the kernel consistently makes poor paging choices and there is no way to force it to make good paging choices. You have much better visibility into the I/O behavior of your application than the kernel does. In my experience, at least for database-y workloads, if you care enough about paging behavior to use madvise(), you might as well just use any number of O_DIRECT alternatives that offer deterministic paging control. It is much easier than trying to cajole the kernel into doing what you need via madvise().
- silon42 6y agommap has less deterministic memory pressure and more complex interactions with overcommit (if enabled).
- __turbobrew__ 6y agoI write software for a HPC centre and we noticed this as well. The speed of our programs heavily using memory maps can vary by an order of magnitude based upon the memory pressure on the node. This non-determinism in runtime is a show stopper for us so we ripped out most memory maps from our codebase. There is just too much magic happening under the hood to have control over your program.
- btown 6y agoCurious now - were you running an unconventional workload that stressed LevelDB, or do you think some version of this advice could be applicable to typical workloads?
- jstimpfle 6y agoYup, I don't like using mmap() for the reason alone that it means giving up a lot of control.
- jeffbee 6y agoLevelDB is kinda like a single-tablet bigtable, but because of that its mmap i/o is not a result of battle hardening in production systems. bigtable doesn't use local unix i/o for any purpose at all, so I'm not surprised to hear that leveldb's local i/o subsystem is half baked.
- BikiniPrince 6y agoOne mechanism we developed was to build a variant of our storage node that could run in isolation. This meant that synthetic testing would give us some optimal numbers for hardware vetting and performance changes. I proved quite quickly our application was quite thread poor and the costs of fixing it was quite worth it. Using other synthetic benchmarks to compare what the systems were capable of. I was gone before that was finished, but it was quite an improvement. It also allowed cold volumes to exist in an over subscription model. None of this excuses good real world telemetry and evaluation of your outliers.