4 ms·
Thanks! Have you read: Can Applications Recover from fsync Failures? — https://www.usenix.org/system/files/atc20-rebello.pdf https://www.usenix.org/system/files
by _vvhw 5y ago
Thanks! Have you read: Can Applications Recover from fsync Failures? — https://www.usenix.org/system/files/atc20-rebello.pdf https://www.usenix.org/system/files/atc20-rebello.pdf
The paper shows that for databases requiring durability and redo logging, O_DIRECT is a good idea for safety.
I do enjoy working with O_DIRECT. I find it works best when designing systems from the ground up for O_DIRECT. It leads to a very simple system in the end that works equally well on block devices, which is a good place to be. And the performance gains by not thrashing the CPU cache through memcopies to the page cache are nice.
- ignoramous 5y agoAll db systems (even and especially the replicated ones) at AWS had to use O_DIRECT (I presume for this same reason).
- _vvhw 5y agoIt's interesting how the proper handling of local storage faults is now also recognized to be all the more critical for global replicated systems—that local faults do propagate across distributed systems. For example, when not using O_DIRECT, an ack back to the consensus protocol, for local data that was recovered from the log at startup, and not in fact made durable (only marked clean in the kernel page cache after an fsync failure), could cause a quorum swing after the next reboot, and global data loss.