38 ms·
Are You Sure You Want to Use MMAP in Your Database Management System? (2022)
- Dwedit 3y agoMemory-Mapped Files = access violations when a disk read fails. If you're not prepared to handle those, don't use memory-mapped files. (Access violation exceptions are the same thing that happens when you attempt to read a null pointer) Then there's the part with writes being delayed. Be prepared to deal with blocks not necessarily updating to disk in the order they were written to, and 10 seconds after the fact. This can make power failures cause inconsistencies.
- kentonv 3y ago> Be prepared to deal with blocks not necessarily updating to disk in the order they were written to, and 10 seconds after the fact. This can make power failures cause inconsistencies. This is not specific to mmap -- regular old write() calls have the same behavior. You need to fsync() (or, with mmap, msync()) to guarantee data is on disk.
- crabbone 3y ago> This is not specific to mmap -- regular old write() calls have the same behavior. This is not true. This depends on how the file was opened. You may request DIRECT | SYNC when opening and the writes are acknowledged when they are actually written. This is obviously a lot slower than writing to cache, but this is the way for "simple" user-space applications to implement their own cache. In the world of today, you are very rarely writing to something that's not network attached, and depending on your appliance, the meaning of acknowledgement from write() differs. Sometimes it's even configurable. This is why databases also offer various modes of synchronization -- you need to know how your appliance works and configure the database accordingly.
- kentonv 3y ago> This is not true. This depends on how the file was opened. You may request DIRECT | SYNC Well sure, but 99.9% of people don't do that (and shouldn't, unless they really know what they are doing). > In the world of today, you are very rarely writing to something that's not network attached, and depending on your appliance, the meaning of acknowledgement from write() differs. What network-attached storage actually uses O_SYNC behavior without being asked? I'd be quite surprised if any did this as it would make typical workloads incredibly slow in order to provide a guarantee they didn't ask for.
- pclmulqdq 3y ago100% of people writing a database know about filesystem options like DIRECT and SYNC, and that is the subject of this paper. Also, most of the network-attached storage we people use is in the form of things like EBS, which is very careful to imitate the behavior of a real disk, but with different performance and some different (albeit very rare) failure modes.
- tsimionescu 3y agoIt's fun to remember that fsync() on Linux on ext4 at least offers no real guarantee that the data was successfully written to disk. This happens when write errors from background buffered writes are handled internally by the kernel, and they cleanup the error situation (mark dirty pages clean etc). Since the kernel can't know if a later call to fsync() will ever happen, it can't just keep the error around. So, when the call does happen, it will not return any error code. I don't know for sure, but msync() may well have the same behavior. Here is an LWN article discussing the whole problem as the Postgres team found out about it. https://lwn.net/Articles/752063/ https://lwn.net/Articles/752063/
- sidewndr46 3y agodoes that get delivered as SIGSEGV to the process or something else?
- afr0ck 3y agoOn Linux, it's a SIGBUS.
- afr0ck 3y agoLinux throws a SIGBUS. A process should anticipate such I/O failures by implementing a SIGBUS handler, especially a database server. For the second part of your comment, on Linux systems, there is the msync() system call that can be used to flush the page cache on demand.
- crabbone 3y ago> msync() system call that can be used to flush the page cache on demand. for everyone, not just the file you mapped to memory. I.e. the guarantee is that your file will be written, but there's no way to do that w/o affecting others. This is not such a hot idea in an environment where multiple threads / processes are doing I/O.
- afr0ck 3y agomsync() affects only the pages that part of the mmap area you ask for in the arguments. From the man pages: > int msync(void addr[.length], size_t length, int flags); > msync() flushes changes made to the in-core copy of a file that was mapped into memory using mmap(2) back to the filesystem
- crabbone 3y agoNo it doesn't. That's physically impossible. Read what you quoted -- it never says that it's going to do it only for the file in question. If you don't know why it's not possible, here's a simplified version of it: hardware protocols (s.a. SCSI) must have fixed size messages to fit them through the pipeline. I.e. you cannot have a message larger than the memory segment used for communication with the device, because that will cause fragmentation and will lead to a possibility of message being corrupted (the "tail" being lost or arriving out of order). On the other hand, to "flush" a file to persistent storage you'd have to specify all blocks associated with the file that need to be written. If you try to do this, it will create a message of arbitrary size, possibly larger than the memory you can store it in. So, the only way to "flush" all blocks associated with a file is to "flush" everything on a particular disk / disks used by the filesystem. And this is what happens in reality when you do any of the sync family commands. The difference is only in what portion of the not-yet synced data the OS will send to the disk before requesting a sync, but the sync itself is for the entire disk, there aren't any other syncs.
- wmf 3y agoI wonder how many apps don't handle errors from read() anyway.
- zffr 3y agoThe TLDR is that MMAP sorta does what you want, but DBMSes need more control over how/when data is paged in/out of memory. Without this extra control, there can be issues with transactional safety, and performance.
- jasonhansel 3y agoI've become convinced that there are very few, if any, reasons to MMAP a file on disk. It seems to simplify things in the common case, but in the end it adds a massive amount of unnecessary complexity.
- AnotherGoodName 3y agoComplexity? You mmap it in and then read the multi terrabyte file as if it was an array. The opposite with actual file io sucks in terms of complexity. I get that you can write bespoke code that performs better but mmap is a one liner to turn a file into an array.
- Dwedit 3y agoNeed to handle the exceptions/signals every time a disk read fails. With classic IO, you know when the read will happen. But with memory-mapped files, the exception can happen at any time you are reading from the memory range. As for why disk reads fail, yes that's a thing. Less common on internal storage (bad sectors), but more common on removable USB devices or Network drives (especially on wifi).
- Sesse__ 3y agoMulti-terabyte? Better hope you have lots of spare RAM for all those page structures the kernel has to keep.
- vvanders 3y agoIt's incredibly useful in read-only, memory constrained scenarios. I.E. we used to mmap all of our animation data on many rendering engines I worked on where having ~20-50mb of animation data and only "paying" a couple 10s of kb based on usage patterns was very handy. It becomes even more powerful when you have multiple processes sharing that data and the kernel is able to re-use clean pages across processes. From reading the paper most of the concerns are around the write side. LMDB is the primary implementation that I know which leans heavily into mmap but it also comes with a number of constraints there(single writer, read locks can lead to unbounded appending to the WAL, etc). As with any tech choice it's about knowing constraints/trade-offs and making appropriate choices for your domain.
- jFriedensreich 3y agomaybe a stupid question but what is wrong with coffee and spicy food?
- toxik 3y agoJust doesn’t taste good together I think
- orf 3y agoFor the majority of the world, nothing. But if your diet consists of fairly bland food then it can result in unpleasant trips to the toilet.
- mattnewton 3y agoAcid reflux I thought
- pizza 3y agoto put it crudely I think the punchline is the spicy food hurts on the way out, and the coffee makes that happen with greater velocity
- wood_spirit 3y agoOld timers will recall when using mmap was a prominently promoted selling point for the “no sql” dbms.
- nemo44x 3y agoFor documents it made access fast since there’s no joins, etc. that require paging from all over. The problem ended up being updates and compaction issues.
- wood_spirit 3y agoMy memory is that the problem was ACID. The document stores didn’t promise to be reliable because apparently that didn’t scale. And there was a very well known cartoon video discussion about it with “web scale” and “just write to dev null” and other classics that became memes :)
- cratermoon 3y agoDid you ever read Pat Helland's article, "Life Beyond Distributed Transactions: An apostate’s opinion" https://dl.acm.org/doi/10.1145/3012426.3025012 https://dl.acm.org/doi/10.1145/3012426.3025012? "This article explores and names some of the practical approaches used in the implementation of large-scale mission-critical applications in a world that rejects distributed transactions."
- wood_spirit 3y agoNo I haven’t. Thanks for the interesting link :) Admittedly I live in a world where big distributed transactions are a given and work fine and sql speeds us up not slows us down. I’m guessing sql and acid scaled after all?
- cratermoon 3y ago> I’m guessing sql and acid scaled after all? Yes and no. Distributed transactions and two-phase commit have been superseded by things like Paxos and Raft, with a variety of consistency models, so the implementation is drastically different.
- dist1ll 3y agoMany general-purpose OS abstractions start leaking when you're working on systems-like software. You notice it when web servers are doing kernel bypass to for zero-copy, low-latency networking, or database engines throw away the kernel's page cache to implement their own file buffer.
- arter4 3y agoWeb servers doing kernel bypass for zero-copy networking? Do you have a specific example in mind? I'm curious.
- kentonv 3y agoProbably the most common example is sendfile() for writing file contents out to a socket without reading them into userspace: https://man7.org/linux/man-pages/man2/sendfile.2.html https://man7.org/linux/man-pages/man2/sendfile.2.html
- arter4 3y agoYes, I knew about sendfile() but I wasnt't aware of any web server using that (though I know Kafka uses it). Then I found out Apache supports it via the EnableSendfile directive. Nice. >This directive controls whether httpd may use the sendfile support from the kernel to transmit file contents to the client. By default, when the handling of a request requires no access to the data within a file -- for example, when delivering a static file -- Apache httpd uses sendfile to deliver the file contents without ever reading the file if the OS supports it.
- AnotherGoodName 3y agoA well written bespoke function can beat a generalized function at a specific task. If you have the resources to write and maintain the bespoke method great. The large database developers probably have this. For others please don't take this link and go around claiming mmap is bad though. That gets tiresome and is misguided. Mmap is a shortcut to access large files in a non linear fashion. It's good at that too. Just not as good as a bespoke function.
- formerly_proven 3y agommap can be handy but usually is not a good idea when you care about ACID properties. So it tends to be most useful outside databases.
- josephg 3y agoCan you give some examples where mmap is useful?
- deleted 3y ago[deleted]
- dataflow 3y agoIf your data is likely to already be in the system cache, memory mapping can achieve zero copying of the data, whereas reading will perform at least one memcpy. So there can be a performance advantage depending on the usage pattern. Also, I've never tested this, but I believe mapped files will get flushed as long as the system stays running. So if you only need resilience against abnormal termination rather than system crashes, it seems like a good option?
- amluto 3y ago> Also, I've never tested this, but I believe mapped files will get flushed as long as the system stays running. So if you only need resilience against abnormal termination rather than system crashes, it seems like a good option? Linux will not lose data written to a MAP_SHARED mapping when the process crashes. But! Linux will synchronously update mtime when starting to write to a currently write protected mapping (e.g. one which was just written out). This means (a) POSIX is violated (IMO) and (b) what should be a minor fault to enable writes turns into an actual metadata write, which can cause actual synchronous IO. I have an ancient patch set to fix this, but I never got it all the way into upstream Linux. What you can do is mmap a file on a tmpfs as long as you trust yourself to have some other reliable process handle the data even if your application terminates abnormally. This is awkward with a container solution if you need to survive termination of the entire container.
- SoftTalker 3y agoThis reads more like "don't write your own DBMS" than "don't use mmap."
- jandrewrogers 3y agoAnother interesting limitation of mmap() is that real-world storage volumes can exceed the virtual address space a CPU can address. A 64-bit CPU may have 64-bit pointers but typically cannot address anywhere close to 64 bits of memory, virtually or physically. A normal buffer pool does not have this limitation. You can get EC2 instances on AWS with more direct-attached storage than addressable virtual address space on the local microarchitecture.
- glandium 3y agoTo put concrete numbers: x86-64 is limited to 48 bits for virtual addresses, which is "only" 256TiB (281TB).
- stevefan1999 3y agoIntel now extended the page table level to 5-level making this number not so valid. Granted, PL5 creates more TLB pressure and longer memory access time due to that.
- Svetlitski 3y agoStarting with Ice Lake there’s support for 5-level paging, which increases this to 128 PiB. Can’t say that I’ve ever seen this used in the wild though.
- jandrewrogers 3y agoYeah, there mostly isn’t a use case for it in databases. If you have that much storage you’ll need to bypass the kernel cache and scheduler anyway for other reasons. That was true even at the 48-bit limit.
- hyc_symas 3y agoAll of that is true, but I don't think it's a realistic concern. You're going to be sharding your data across multiple nodes before it gets that large. Nobody wants to sit around backing up or restoring a monolithic 256 TiB database.
- deleted 3y ago[deleted]
- kwohlfahrt 3y agoIt sounds like a lot of the performance issues are TLB-related. Am I right in thinking huge-pages would help here? If so, it's a bit unfortunate they didn't test this in the paper. Edit: Hm, it might not be possible to mmap files with huge-pages. This LWN article[1] from 5 years ago talks about the work that would be required, but I haven't seen any follow-ups. [1]: https://lwn.net/Articles/718102/ https://lwn.net/Articles/718102/
- ori_b 3y agoNo, huge pages wouldn't help. They would change when the TLB gets flushed, but the flushes would still be there.
- hyc_symas 3y agoHuge pages aren't pageable, so they wouldn't be particularly advantageous for a mmap DB anyway, you'd have to do traditional I/O & buffer management for everything.
- pjdesno 3y agoNot just databases - we ran into the same issues when we needed a high-performance caching HTTP reverse proxy for a research project. We were just going to drop in Varnish, which is mmap-based, but performance sucked and we had to write our own. Note that Varnish dates to 2006, in the days of hard disk drives, SCSI, and 2-core server CPUs. Mmap might well have been as good or even better than I/O back then - a lot of the issues discussed in this paper (TLB shootdown overhead, single flush thread) get much worse as the core count increases.
- Sesse__ 3y agoVarnish' design wasn't very fast even for 2006-era hardware. It _was_ fast compared to Squid, though (which was the only real competitor at the time), and most importantly, much more flexible for the origin server case. But it came from a culture of “the FreeBSD kernel is so awesome that the best thing userspace can do is to offload as many decisions as humanly possible to the kernel”, which caused, well, suboptimal performance. AFAIK the persistent backend was dropped pretty early on (eventually replaced with a more traditional read()/write()-based one as part of Varnish Plus), and the general recommendation became just to use malloc and hope you didn't swap.
- tayo42 3y agoVarnish has a file system backed cache that depends on the page cache to keep it fast. What did you differently in your custom one that was faster then varnish?
- pjdesno 3y agoSimple multithreaded read/write. On a 20-core 40-thread machine with a couple of fast NVMe drives it was way faster.
- hyc_symas 3y agoThis is a pretty old argument and IMO it's far out of date/obsolete. Taking full control of your I/O and buffer management is great if (a) your developers are all smart and experienced enough to be kernel programmers and (b) your DBMS is the only process running on a machine. In practice, (a) is never true, and (b) is no longer true because everyone is running apps inside containers inside shared VMs. In the modern application/server environment, no user level process has accurate information about the total state of the machine, only the kernel (or hypervisor) does and it's an exercise in futility to try to manage paging etc at the user level. As Dr. Michael Stonebraker put it: The Traditional RDBMS Wisdom is (Almost Certainly) All Wrong. https://slideshot.epfl.ch/play/suri_stonebraker https://slideshot.epfl.ch/play/suri_stonebraker (See the slide at 21:25 into the video). Modern DBMSs spend 96% of their time managing buffers and locks, and only 4% doing actual useful work for the caller. Granted, even using mmap you still need to know wtf you're doing. MongoDB's original mmap backing store was a poster child for Doing It Wrong, getting all of the reliability problems and none of the performance benefits. LMDB is an example of doing it right: perfect crash-proof reliability, and perfect linear read scalability across arbitrarily many CPUs with zero-copy reads and no wasted effort, and a hot code path that fits into a CPU's 32KB L1 instruction cache.
- danappelxx 3y agoWho is deploying databases in containers?
- huahaiy 3y agoEmbedded DB
- orbz 3y agoA disturbingly large number of deployments I’ve seen using Kubernetes or docker compose have databases deployed as such.
- danappelxx 3y agoIMO if you’re concerned about performance and yet are deploying databases this way — mmap should not even be on the radar.
- mpweiher 3y agoYes, I definitely would want to use mmap() in my storage system. And would love to see the limitations that make this tricky addressed.
- benlivengood 3y agoFor all of its usefulness in the good old days of rusty disks I wonder if virtual memory is worth having for dedicated databases, caches, and storage heads. Avoiding TLB flushes entirely sounds like a huge win for massively multithreaded software and memory management in a large shared flat address space doesn't sound impossibly hard.
- hyc_symas 3y agoThe jump in address sizes starts to get too unwieldy. 32 bit addresses were ok, 64 bit addresses start to get clunky, 128 bit would be exorbitant for CPU real estate. There's a reason AMD64 still only supported 40 physical address bits when it was introduced, and later only expanded to 48 bits. The reality is there will always be a hierarchy for storage, and paging will always be the best mechanism to deal with it. Because primary memory will always be most expensive, no matter what technology it's based on. There will always be something slower, cheaper, and denser that will be used for secondary storage. There will always be cheaper storage. And its capacity will exceed primary, and it will always be most efficient to reference secondary storage in chunks - pages - and not at individual byte addresses.
- moonchild 3y agoI don't really see what those two things have to do with each other. When you don't use mmap, you manage the disc<->ram storage virtualisation yourself. Hardware paging, then, is pure overhead. The parent doesn't argue against layering of storage media, nor against chunking in general. Only against mmus as a mechanism for implementing it.
- dang 3y agoRelated: Are You Sure You Want to Use MMAP in Your Database Management System? [pdf] - https://news.ycombinator.com/item?id=31504052 https://news.ycombinator.com/item?id=31504052 - May 2022 (43 comments) Are you sure you want to use MMAP in your database management system? [pdf] - https://news.ycombinator.com/item?id=29936104 https://news.ycombinator.com/item?id=29936104 - Jan 2022 (127 comments)