5 ms·
I am going to write this comment with a large preface: I don't think it is ever helpful to be an absolutist. For every best-practice/"right way" to do things, t
by addisonj 3y ago
I am going to write this comment with a large preface: I don't think it is ever helpful to be an absolutist. For every best-practice/"right way" to do things, there are circumstances when doing it another way makes sense. That can be a ton of reasons for that, be it technical, money/time, etc. The best engineering teams aren't those that just blindly follow what others say is a best practice but understand the options and make an informed choice. None of the following comment is at all commentary on questDB, as they mention in the article, many databases use similar tools.
With that said, after reading the first paragraph I immediately searched the article for "mmap" and had a good sense of where the rest of this was going. Put simply, it is just really hard to consider what the OS is going to do in all situations when using mmap. Based on my experience, I would guess that a ton of people reading this comment have hit issues that, I would argue, is due to using mmap. (Particularly looking at you prometheus).
All things told, this is a pretty innocuous incident of mmap causing problems, but I would encourage any aspiring DB engineers to read https://db.cs.cmu.edu/mmap-cidr2022 https://db.cs.cmu.edu/mmap-cidr2022 as it gives a great overview of the range of problems that can occur when using mmap
I think some would argue that mmap is "fine" for append only workloads (and is certainly more reasonable compared to a DB with arbitrary updates) but even here, lots of factors like metadata, scaling number of tables, etc will eventually bring you to hit some fundamental problems when using mmap.
The interesting opportunity in my mind, especially with improvements in async IO (both at FS level and in tools like rust), is to build higher level abstractions that bring the "simplicity" of mmap, but with more purpose-built semantics ideal for databases.
- Sytten 3y agoWhen I read they were using mmap I immediately thought of Andy Pavlo since he warns against it every time he can in the CMU videos. Enough that I thought it was somewhat of a concensus now that mmap should be avoided specially for databases, guess I was wrong.
- eatonphil 3y ago> a concensus now that mmap should be avoided specially for databases Maybe, but table 1 in Andy Pavlo's paper shows 7 of 10 databases surveyed do still use mmap. Furthermore, that paper clearly demonstrating the issues with mmap came out in 2022 and most databases have been around longer than that. That mmap isn't the future is maybe more certain than that mmap isn't common practice today (because it does sorta seem to be).
- loeg 3y agoMy impression is that databases have known of the shortcomings of mmap since, like, the 90s. A critical flaw is the lack of error handling -- in addition to the unpredictable performance characteristics and naive caching. I'm looking forward to reading the paper.
- josephg 3y ago> Enough that I thought it was somewhat of a concensus now that mmap should be avoided specially for databases I have to remind myself regularly that the world is larger than it seems. The other day someone on r/rust said "Surely everyone knows what the rust programming language is by now. Can we stop introducing rust every time we mention it in a paper?". But no, obviously everyone in your circles knows what rust is. But you don't know most developers. And I wouldn't be surprised if less than half of working developers have heard of rust, even if everyone you know knows about it. I used to rant against using socket.io at every possible opportunity. The library has (had) crazy bugs in its reconnection code. In the right circumstances the library would violate ordering and delivery guarantees, or it would lie about messages being received when they hadn't been. But no matter how much I ranted about it, and no matter how many hundreds of issues there were on github, far more people used socket.io than the (much more reliable) alternatives because socket.io had a pretty website, good documentation and it was taught at coding bootcamps. I think the only reason its not as popular now is that you don't need it now that websockets are available everywhere. My partner says she imagines asking questions of our families when she tries to imagine what the average person thinks. But our immediate families are still a really weird bubble - every single one of the adults has graduated from college. (And weirdly, over half of that group have also taught at college.) That's still a really biased set of people. Finding an unbiased set is wildly difficult.
- postdb 3y ago>I used to rant against using socket.io at every possible opportunity. The library has (had) crazy bugs in its reconnection code. In the right circumstances the library would violate ordering and delivery guarantees, or it would lie about messages being received when they hadn't been. But no matter how much I ranted about it, and no matter how many hundreds of issues there were on github, far more people used socket.io than the (much more reliable) alternatives because socket.io had a pretty website, good documentation and it was taught at coding bootcamps. I think the only reason its not as popular now is that you don't need it now that websockets are available everywhere. That is true. So with the modern browsers now days latest Chrome/Firefox (both desktop and mobile) can support websocket seamlessly? I guess socket.io is kinda like the jquery of ws then? It will take some time to phase out.
- Sesse__ 3y agommap should generally be avoided, not just for databases. It's useful for quick prototyping and for the specific case of demand-paging executables (which is really what it's made for!), but there are so many pitfalls overall. You can't mmap large files relative to your memory (you crash into either address space limits or PTE memory usage), you have absolutely no hope of recovering from errors, you can't do large sequential I/O reliably, it's a really difficult problem to order your writes, and so on. There are corner cases where it's great, like when you have a file that you know is 90% in-core already and you don't care about errors. But overall, read() and write() are simpler, faster, more reliable primitives.
- citrin_ru 3y agoHere is a good use case for mmap - a process-a performs data processing and writes results to a disk (or tmpfs), next you need to repeatedly read in a process-b (or read once but not sequentially). If you'll use read() you will: 1. double RAM usage - the file will be in a VM cache anyway (unless you'll use O_DIRECT which would make this pipeline slower) and without mmap() you'll have to create 2nd copy inside the process-b 2. add unnecessary kernel->userspace copy while reading data in process-b. But for saving data from process-a I would still use write() using MAXPHYS sized blocks: I'm not sure mmap would use optimal write block size, and with write it is easier detect errors (like ENOSPC).
- Sesse__ 3y agoThat's a good use case for… a quite regular pipe? :-)
- citrin_ru 3y agoPipe would not allow receiving process to read the data more than one time or jump across the data from a location to a location (without storing either in memory or on FS). Another good use case for mmap is sharing read-only (or rarely updated) dataset among multiple processes.
- m463 3y agoSo a dbms probable has intimate knowledge about its data that probably can't be hinted via madvise. But I wonder when there are decent reasons to let the OS handle file I/O through the demand paging system, since it's good at it.
- loeg 3y ago> since it's good at it. This is generous. At least, it's not a great working assumption. If you know anything about your workload, it's often possible to do better by specializing slightly.
- justin66 3y ago> All things told, this is a pretty innocuous incident of mmap causing problems, but I would encourage any aspiring DB engineers to read https://db.cs.cmu.edu/mmap-cidr2022 https://db.cs.cmu.edu/mmap-cidr2022 as it gives a great overview of the range of problems that can occur when using mmap At first glance (doing a few text finds and a really quick read through the paper after clicking through that intro page with the poop emoji at the top and disregarding Recommended Music for this Paper: Dr. Dre – High Powered (featuring RBX)) that paper seems too short to adequately explore the topic. On the other hand these same guys (the CMU Database Group) are an amazing resource and their youtube channel offers some great stuff [1] that would allow a curious person to explore the topic in greater depth if they dug into the papers of everyone who gave presentations at CMU. clickbait: How many of the world's leading software engineers who addressed CMU students and professors about their successful database products rely on mmap? The answer may surprise you. [1] https://www.youtube.com/@CMUDatabaseGroup https://www.youtube.com/@CMUDatabaseGroup
- ayende 3y agoI wrote a response to this article, because that is a bad comparison. https://ravendb.net/articles/re-are-you-sure-you-want-to-use-mmap-in-your-database-management-system https://ravendb.net/articles/re-are-you-sure-you-want-to-use... Agree on CMU being a great resource.
- justin66 3y agoThanks. I didn't know about this, but I read the stuff you wrote about lmdb back in the day.
- apavlo 3y ago> disregarding Recommended Music for this Paper: Dr. Dre – High Powered (featuring RBX)) Why? > that paper seems too short to adequately explore the topic. The paper was published in CIDR (https://www.cidrdb.org https://www.cidrdb.org). The paper submissions for this conference are meant to be short (typically 6-7 pages) to ensure that people can get their ideas out quickly.