5 ms·
O_DSYNC on Darwin is completely broken for the last 15 years
- ayende 4y agoThat is a killer for performance for databases
- pengaru 4y ago> That is a killer for performance for databases It's actually the thing you use instead of O_SYNC because it's expected to deliver better synchronous performance, tantamount to using fdatasync() vs. fsync(). A broken fdatasync() (or O_DSYNC) is kind of a serious problem.
- _vvhw 4y agoWe discovered this through our use of O_DSYNC, and O_SYNC also has the same problem. I think that for macOS so far, the focus has been on ordering guarantees, i.e. write barriers to reset the disk back to some atomic point in time after a crash. However, for distributed databases like TigerBeetle, we also need durability guarantees, i.e. be able to know that the disk has the data down cold, that the disk won't “forget” data that the node has underwritten and ACK'ed externally.
- _vvhw 4y agoYes, we were counting on O_DSYNC to set the FUA bit on the write syscall so that we could avoid an extra F_FULLFSYNC fcntl syscall.
- qhwudbebd 4y agoTwo other macos breakages that I've recently been bitten by: They never got round to implementing unnamed posix semaphores: sem_init() will always give EINVAL. They never got round to finishing poll(). It fails on devices. Because this includes /dev/null, it can make using poll() on stdin and stdout in general-purpose command-line tools dangerous. It used to fail on pipes too, but possibly this is now fixed?
- VogonPoetry 4y agoOpinion. The bizarro thing with this is that O_<xzy> flags are for a per file open, but the guarantees can only be supported by an underlying filesystem. On Linux you can mount a filesystem and invalidate some metadata updates using options such as noatime. This is also depends on the mounted filesystem, such as FAT, which might not even support the updated attributes. Are these constraints honored in these environments with these flags? (sorry, too lazy to write code and check this) There is no data on what is actually "broken" or why it might violate POSIX.1-2008. What metadata is actually important or impactful? I am very, very skeptical of claims made by some database folks, due to past experience. Some Oracle products, and from what I understand, still do not understand that monotonically increasing means greater than or equal to (>=). The equal to is the problem, because calling an absolute value timestamp function to derive a "unique" value will always fail (with the same return value) when you get faster hardware and more cores. In a past life (~2000), I had to add code to bork an operating system to appease this very misguided notion. Last month (May 2022) I got (2nd hand) info that after a hardware update to faster hardware, duplicate unique row_ids were being created in an Oracle DB. Without additional details, I find it hard to understand what the issue actually is. Is this a detectable temporal violation of order of operations for a live system, or some sort of fail situation where an abort / reboot occurs -- would love to understand how the external logging of this was carried out, because how was that synchronized / timestamped?
- _vvhw 4y agoDetails and verification are in the quoted thread (you need to click through). TLDR: https://twitter.com/jorandirkgreef/status/1532314169604726784 https://twitter.com/jorandirkgreef/status/153231416960472678... You can see that, for those block sizes, the throughput of O_DSYNC is the same as F_NOCACHE, e.g. 3121 MiB/s vs 52.05 MiB/s for libuv's durable fdatasync() which is really actually using fcntl(F_FULLFSYNC) under the hood here. In other words, O_DSYNC is only going as far as the disk's own internal cache. It's giving nearly the same performance as a buffered write. So the reason it is “completely broken” and not simply “broken” then, is that setting the O_DSYNC bit provides no durability at all. It's a no-op. Also, to be fair, if you're familiar with the space, it's been known for some time that Apple's fsync() is not durable—it's not like this is coming out of left field. What's new here, is that O_SYNC/O_DSYNC have the same issue. That's why we're awarding the bounty.