10 ms·
But how, exactly, do databases use mmap?
- PaulHoule 6y agoI like mmap and I don't. It is incompatible with non-blocking I/O since your process will be stopped if it tries to access part of the file that is not mapped -- this isnt a syscall blocking (which you might work around) but rather any attempt to access mapped memory. I like mmap for tasks like seeking into ZIP files, where you can look at the back 1% of the file, then locate and extract one of the subfiles; the trouble there is that the really fun case is to do this over the network with http (say to solve Python dependencies, to extract the metadata from wheel files) in which case this method doesnt work.
- quotemstr 6y agoYou use mmap whether you want to or not: the system executes your program by mmaping your executable and jumping into it! You can always take a hard fault at any time because the kernel is allowed to evict your code pages on demand even if you studiously avoid mmap for your data files. And it can do this eviction even if you have swap turned off. If you want to guarantee that your program doesn't block, you need to use mlockall.
- jorangreef 6y agoBut that's a different order of magnitude problem: control plane vs data plane. At some point, we could also say that the line fill buffer blocks our programs (more often than we realize). All of this is accurate, but at different scales.
- PaulHoule 6y agoAlso many systems in 2021 have a lot of RAM and hardly ever swap.
- geofft 6y agoThis is technically true, but the use case we're talking about is programs that are much smaller than their data. Postgres, for instance, is under 50 MB, but is often used to handles databases in the gigabytes or terabytes range. You can mlockall() the binary if you want, but you probably can't actually fit the entire database into RAM even if you wanted to. Also, when processing a large data file (say you're walking a B-tree or even just doing a search on an unindexed field), the code you're running tends to be a small loop, within the same few pages, so it might not even leave the CPU's cache, let alone get swapped out of RAM, but you need to access a very large amount of data, so it's much more likely the data you want could be swapped out. If you know some things about the data structure (e.g., there's an index or lookup table somewhere you care about, but you're traversing each node once), you can use that to optimize which things are flushed from your cache and which aren't.
- quotemstr 6y agoIndeed. It's a question of scale: I write programs that can't afford to get blocked behind IO, ever, and that level, I need to pay attention to things like code paging, and even more esoteric things like synchronous reclaim. If you're just optimizing stuff generally instead of trying to guarantee invariants, sure, ignore code paging and use direct IO for your own data.
- loeg 6y agoYou're not wrong. Applications and libraries that want to be non-blocking should mlock their pages and avoid mmap for further data access. ntpd does this, for example. After application startup, you can avoid additional mmap.
- rapsey 6y agoProcess will be stopped or thread?
- ithkuil 6y agoThread
- amelius 6y ago> It is incompatible with non-blocking I/O since your process will be stopped if it tries to access part of the file that is not mapped Yeah, but the same problem occurs in normal memory when the OS has swapped out the page. So perhaps non-blocking I/O (and cooperative multitasking) is the problem here.
- loeg 6y ago> Yeah, but the same problem occurs in normal memory when the OS has swapped out the page. I'd argue that swapping is an orthogonal problem which can be solved in a number of ways: disable swap at the OS level, mlock() in the application, maybe others. mmap is really a bad API for IO — it hides synchronous IO and doesn't produce useful error statuses at access. > So perhaps non-blocking I/O (and cooperative multitasking) is the problem here. I'm not sure how non-blocking IO is "the problem." It's something Windows has had forever, and unix-y platforms have wanted for quite a long time. (Long history of poll, epoll, kqueue, aio, and now io_uring.)
- amelius 6y ago> it hides synchronous IO and doesn't produce useful error statuses at access. You can trap IO errors if necessary. E.g. you can raise signals just like segfaults generate signals. > I'm not sure how non-blocking IO is "the problem." The point is that non-blocking IO wants to abstract away the hardware, but the abstraction is leaky. Most programs which use non-blocking IO actualy want to implement multitasking without relying threads. But that turns out to be the wrong approach.
- loeg 6y ago> The point is that non-blocking IO wants to abstract away the hardware, but the abstraction is leaky. Why do you say it doesn't match hardware? Basically all hardware is asynchronous — submit a request, get a completion interrupt, completion context has some success or failure status. Non-blocking IO is fundamentally a good fit for hardware. It's blocking IO that is a poor abstraction for hardware. > Most programs which use non-blocking IO actualy want to implement multitasking without relying threads. But that turns out to be the wrong approach. Why is that the wrong approach? Approximately every high-performance httpd for the last decade or two has used a multitasking, non-blocking network IO model rather than thread-per-request. The overhead of threads is just very high. They would like to use the same model for non-network IO, but Unix and unix-alikes have historically not exposed non-blocking disk IO to applications. io_uring is a step towards a unified non-blocking IO interface for applications, and also very similar to how the operating system interacts with most high-performance devices (i.e., a bunch of queues).
- Sesse__ 6y agommap is great for rapid prototyping. For anything I/O-heavy, it's a mess. You have zero control over how large your I/Os are (you're very much at the mercy of heuristics that are optimized for loading executables), readahead is spotty at best (practical madvise implementation is a mess), async I/O doesn't exist, you can't interleave compression in the page cache, there's no way of handling errors (I/O error = SIGBUS/SIGSEGV), and write ordering is largely inaccessible. Also, you get issues such as page table overhead for very large files, and address space limitations for 32-bit systems. In short, it's a solution that looks so enticing at first, but rapidly costs much more than it's worth. As systems grow more complex, they almost inevitably have to throw out mmap.
- codetrotter 6y ago> the trouble there is that the really fun case is to do this over the network with http (say to solve Python dependencies, to extract the metadata from wheel files) in which case this method doesnt work If the web server can tell you the total size of the file by responding to a HEAD request, and it support range requests then it will be possible. https://developer.mozilla.org/en-US/docs/Web/HTTP/Range_requests https://developer.mozilla.org/en-US/docs/Web/HTTP/Range_requ... Or am I missing something?
- amelius 6y agoThis is one area where Rust, a modern systems language, has disappointed me. You can't allocate data structures inside mmap'ed areas, and expect them to work when you load them again (i.e., the mmap'ed area's base address might have changed). I hope that future languages take this usecase into account.
- quotemstr 6y agoYou can't do that in C++ or any language. You need to do your own relocations and remember enough information to do them. You can't count on any particular virtual address being available on a modern system, not if you want to take advantage of ASLR. The trouble is that we have to mark relocated pages dirty because the kernel isn't smart enough to understand that it can demand fault and relocate on its own. Well, either that, or do the relocation anew on each access.
- secondcoming 6y agoIt works with C++ if you use boost::interprocess. Its data structures use offset_ptr internally rather than assuming every pointer is on the heap.
- quotemstr 6y agoSure. But that counts as "doing your own relocations". Unsafe Rust could do the same, yes?
- whimsicalism 6y agoWhat is being relocated?
- ithkuil 6y agoIf you use offsets instead of pointers you're doing relocations "on the fly"
- 6y ago
- bonzini 6y agoThe right answer is that they shouldn't. A database has much more information than the operating system about what, how and when to cache information. Therefore the database should handle its own I/O caching using O_DIRECT on Linux or the equivalent on Windows or other Unixes. The article at https://www.scylladb.com/2017/10/05/io-access-methods-scylla/ https://www.scylladb.com/2017/10/05/io-access-methods-scylla... is a bit old (2017) but it explains the trade-offs
- quotemstr 6y agoYep. Every mature, high performing, non-embedded database evolves towards getting the underlying operating system out of the way as much as possible.
- jorangreef 6y agoYes, and it's not only about performance, but also safety because O_DIRECT is the only safe way to recover from the journal after fsync failure (when the page cache can no longer be trusted by the database to be coherent with the disk): https://www.usenix.org/system/files/atc20-rebello.pdf https://www.usenix.org/system/files/atc20-rebello.pdf From a safety perspective, O_DIRECT is now table stakes. There's simply no control over the granularity of read/write EIO errors when your syscalls only touch memory and where you have no visibility into background flush errors.
- formerly_proven 6y agoAround four years ago I was working on a transactional data store and ran into these issues that virtually no one tells you how durable I/O is supposed to work. There were very few articles on the internet that went beyond some of the basic stuff (e.g. create file => fsync directory) and perhaps one article explaining what needs to be considered when using sync_file_range. Docs and POSIX were useless. I noticed that there seemed to be inherent problems with I/O error handling when using the page cache, i.e. whenever something that wasn't the app itself caused write I/O you really didn't know any more if all the data got there. Some two years later fsyncgate happened and since then I/O error handling on Linux has finally gotten at least some attention and people seemed to have woken up to the fact that this is a genuinely hard thing to do.
- minitoar 6y agoInterana mmaps the heck out of stuff. I’ve found that relying on the file cache works great. Though our access patterns are admittedly pretty simple.
- jeffbee 6y agoApparently in a way that the author of the article, and probably the authors of bolt, do not really understand.
- perbu 6y agoThe author notices that Bolt doesn't use mmap for writes. The reason is surprisingly simple, once you know how it works. Say you want to overwrite a page at some locations that isn't present in memory. You'd write to it and you'd think that is that. But when this happens the CPU triggers a page fault, the OS steps in and reads the underlying page into memory. It then relinquishes control back to the application. The application then continues to overwrite that page. So for each write that isn't mapped into memory you'll trigger a read. Bad. Early versions of Varnish Cache struggled with this and this was the reason they made a malloc-based backend instead. mmaps are great for reads, but you really shouldn't write through them.
- cma 6y agoIsn't there a way around this? When coding for graphics stuff writing to GPU mapped memory people usually take pains to turn off compiler optimizations that might XOR memory against itself to zero it out or AND it against 0 and cause a read, and other things like that. https://docs.microsoft.com/en-us/windows/win32/api/d3d12/nf-d3d12-id3d12resource-map https://docs.microsoft.com/en-us/windows/win32/api/d3d12/nf-... > Even the following C++ code can read from memory and trigger the performance penalty because the code can expand to the following x86 assembly code. C++ code: Copy *((int*)MappedResource.pData) = 0; x86 assembly code: Copy AND DWORD PTR [EAX],0 > Use the appropriate optimization settings and language constructs to help avoid this performance penalty. For example, you can avoid the xor optimization by using a volatile pointer or by optimizing for code speed instead of code size. I guess mmapped files still may need a read to know whether to do copy on write, where mapped memory for the CPU in that case is specifically marked for upload only and gets something flagged that writes it regardless of if there is a change, but mmap maybe has something similar? (edit: this seems to say nothing similar is possible with mmap on x86 https://stackoverflow.com/questions/31014515/write-only-mapping-a-o-wronly-opened-file-supposed-to-work https://stackoverflow.com/questions/31014515/write-only-mapp... but how does it work for GPUs? Something to do with fixed pci-e support on the cpu (base address register https://en.wikipedia.org/wiki/PCI_configuration_space https://en.wikipedia.org/wiki/PCI_configuration_space)?
- alaties 6y ago
- waynesonfire 6y agoThanks for diving into this DB! I find it interesting that many databases share such similar architectural principles. NIH. It's super fun to build a database so why not. Also, don't beat yourself over how deep you'll be diving into the design. Why apologize for this? Those that want a deep expository would quickly move on.
- 29athrowaway 6y agomalloc is implemented using mmap. You map memory manually when you need very low level control over memory.
- jeffbee 6y ago`malloc` is not one thing. Some mallocs use mmap and others use brk. Some implementations use both.
- kevin_thibedeau 6y agoSome use neither.
- rcgorton 6y agoI found some of the 'sizing' snippets in the example came across as disingenuous: if you KNOW the size of the file, mmap it initially using that without the looping overhead. And you presumably know how much memory you have on a given system. The description (at least as how I read the article) implies bolt is a truly naive implementation of a key/value DB
- ramoz 6y agoPerhaps a part 2 would dive a bit deeper into os caching and hardware (SSDs, their interfaces etc)
- shoo 6y agoSee also: sublime HQ blog about complexities of shipping a desktop application using mmap [1] and corresponding 200+ comment HN thread [2]: > When we implemented the git portion of Sublime Merge, we chose to use mmap for reading git object files. This turned out to be considerably more difficult than we had first thought. Using mmap in desktop applications has some serious caveats [...] > you can rewrite your code to not use memory mapping. Instead of passing around a long lived pointer into a memory mapped file all around the codebase, you can use functions such as pread to copy only the portions of the file that you require into memory. This is less elegant initially than using mmap, but it avoids all the problems you're otherwise going to have. > Through some quick benchmarks for the way Sublime Merge reads git object files, pread was around ⅔ as fast as mmap on linux. In hindsight it's difficult to justify using mmap over pread, but now the beast has been tamed and there's little reason to change any more. [1] https://www.sublimetext.com/blog/articles/use-mmap-with-care https://www.sublimetext.com/blog/articles/use-mmap-with-care [2] https://news.ycombinator.com/item?id=19805675 https://news.ycombinator.com/item?id=19805675
- rossmohax 6y agommap is not as free as people think. VM subsystem is full of inefficient locks. Here is a very good writeup on a problem BBC encountered with Varnish: https://www.bbc.co.uk/blogs/internet/entries/17d22fb8-cea2-49d5-be14-86e7a1dcde04 https://www.bbc.co.uk/blogs/internet/entries/17d22fb8-cea2-4...
- lrossi 6y ago> huge pressure on the virtual memory (VM) subsystem due to extensive dirty page writeback and page steals. The VM subsystem is constantly modifying page table entries (PTEs). This PTE churn results in frequent translation lookaside buffer (TLB) flushes and many inter-processor interrupts (IPIs) to do so. These TLB flushes have a very negative performance hit. Interesting. I was aware of various mmap limitations, but I didn’t think about the TLB changes/flushes, which obviously come with an important overhead.
- boxfire 6y agoVery strange to see few to no references to io_uring here. I guess it's still too new. As I've seen many times before so much complexity is replicated in userspace to reproduce kernel behavior out of mmap or DIO/AIO, in order to break the latency, caching, and prioritization into a micromanaged state tuned for a narrow set of applications... Then applied to database code used in a myriad of applications which violate those assumptions and have their own needs. io_uring can't take over fast enough.
- jandrewrogers 6y agoYour assumption is correct, io_uring is too new, it isn't available in most LTS kernels. Give it a few years. Also, if you already have a competent io_submit/O_DIRECT implementation then there are few material performance benefits to io_uring for databases. It mostly just cleans up the API. This has value from a code design/maintenance standpoint, particularly since io_submit is lacking in the documentation department, but the lack of kernel support in most environments makes it a poor tradeoff at this time.
- jlokier 6y agoIs it not the case that io_submit can block in some filesystems to handle filesystem metadata (such as block allocation), even with O_DIRECT, whereas io_uring never blocks the submitting thread?
- jandrewrogers 6y agoYes, in theory. In practice, the way io_submit() is actually used in most systems today would not have that issue, and it is designed that way for other practical reasons. You'd want to use io_uring in a similar way. Even if you ignore the blocking aspect, file system metadata modification at runtime is an edge case factory. For database-y type software generally, it is increasingly uncommon to even install a file system. You work with the raw block devices directly, virtualized or otherwise.