6 ms·
Details and verification are in the quoted thread (you need to click through). TLDR: https://twitter.com/jorandirkgreef/status/1532314169604726784 https://twit
by _vvhw 4y ago
Details and verification are in the quoted thread (you need to click through).
TLDR: https://twitter.com/jorandirkgreef/status/1532314169604726784 https://twitter.com/jorandirkgreef/status/153231416960472678...
You can see that, for those block sizes, the throughput of O_DSYNC is the same as F_NOCACHE, e.g. 3121 MiB/s vs 52.05 MiB/s for libuv's durable fdatasync() which is really actually using fcntl(F_FULLFSYNC) under the hood here.
In other words, O_DSYNC is only going as far as the disk's own internal cache. It's giving nearly the same performance as a buffered write.
So the reason it is “completely broken” and not simply “broken” then, is that setting the O_DSYNC bit provides no durability at all. It's a no-op.
Also, to be fair, if you're familiar with the space, it's been known for some time that Apple's fsync() is not durable—it's not like this is coming out of left field. What's new here, is that O_SYNC/O_DSYNC have the same issue. That's why we're awarding the bounty.
- VogonPoetry 4y agoWhat you appear to be saying here is that the "test" is based solely on comparing write performance, not on the conformance to what was written, when it was written and what was written when power was pulled and what was read back later. Is that a fair assessment? This test seems to be based on how "understood" technology works and an expected outcome. Any technology that works better or differently than what you expect will fail. Intel has already proposed MRAM -- aka persistent memory. If I am understanding correctly, your tests would claim that MRAM backed disks would fail your tests, they would perform way too fast. These privatives seem to be super specialized and very dependent on what actually happens. The system call you reference, fsync() was introduced in BSD 4.2 -- 1983. There were no journaling filesystems in 1983. I think it is entirely reasonable to apply ~40 years of filesystem research to invalidate past system call practice and call it out as voodoo / blog fodder without real measurements on current filesystems and hardware and to also document if it fails or succeeds when the plug is pulled. I acknowledge that doing this is hard work and time consuming. You are not going to get a 10 minute tweet from this, but really, isn't that the point of a robust filesystem?
- _vvhw 4y agoNo, I don't believe that the phrase “based solely” would be a fair assessment. That there's a $1,024 bounty being awarded should be reason enough to suggest more than “10 minutes” went into this. While the issue was first reported in a tweet, you should know that it came to us from an experienced database engineer who worked on both Google Spanner and FoundationDB. The initial triage was to compare O_DSYNC with F_FULLFSYNC on the same device, because the results should be within the same order of magnitude, relative to the same device—we already knew that APFS doesn't make the same distinction between O_SYNC vs O_DSYNC that Linux does. "Intel has already proposed MRAM -- aka persistent memory. If I am understanding correctly, your tests would claim that MRAM backed disks would fail your tests, they would perform way too fast." Again, I think you're missing that O_DSYNC was compared relative to the same device, with Darwin's custom F_FULLFSYNC as baseline. Finally, to assume good faith, you can also imagine that once we had the triage in hand, the issue would have been confirmed independently as “correct” by a third party in the best position to make that assessment.