7 ms·
Can Applications Recover from Fsync Failures?
- CGamesPlay 4y ago> all three file systems mark pages clean after fsync fails, rendering techniques such as application-level retry ineffective. However, the content in said clean pages varies depending on the file system; ext4 and XFS contain the latest copy in memory while Btrfs reverts to the previous consistent state. Failure reporting is varied across file systems; for example, ext4 data mode does not report an fsync failure immediately in some cases, instead (oddly) failing the subsequent call. Failed updates to some structures (e.g., journal blocks) during fsync reliably lead to file-system unavailability. And finally, other potentially useful behaviors are missing; for example, none of the file systems alert the user to run a file-system checker after the failure. Surely there's some motivations behind these behaviors and it's not a bug that was implemented in all 3 filesystems, right?
- masklinn 4y ago> Surely there's some motivations behind these behaviors The primary motivation is probably that it's an annoying case to handle, pretty hard to test, and very uncommon. It's also the original behaviour (IIRC from the fsyncgate reports, freebsd had added keeping the buffers dirty in the early aught but other bsds had inherited the ur-behaviour of marking them clean). A second motivation I can think of (with more historical relevance) is that I don't think there's a userland API to tell the kernel to discard dirty pages, so if you can't mark the pages clean there's good chances you've leaked them. In our modern world it's common for writeback errors to be transient (e.g. a USB key that's not ready yet or somesuch) but 40 years back I'm not sure it made much sense. Though I guess network drivers were always a thing and could always have issues.
- username223 4y agoI’d go for historically very uncommon. fsync() meant “write RAM buffers to hard drive,” and if that failed, you were in a world of hurt, and should probably shut down while doing the least amount of additional harm. With NFS, the situation probably changed to “keep trying for a bit, but don’t do anything dramatic.”
- eis 4y agoOne example for marking dirty pages as clean after fsync failure that they mention filesystem developers having given them is a USB stick that has been pulled and keeping dirty pages causing memory leaks in that case. But to me that argument falls flat, they should free the cache if the underlying storage got removed and any subsequent read or write should fail.
- masklinn 4y ago> But to me that argument falls flat, they should free the cache if the underlying storage got removed Except none of that is true, the USB storage could be reattached and the write successful. The system has no way to know whether the failure is transient or permanent. > any subsequent read or write should fail. That’s a different solution than keeping the dirty pages. Instead you discard the pages and lock to failure. IIRC that’s what openbsd implemented in the wake of fsyncgate.
- eis 4y ago> Except none of that is true, the USB storage could be reattached and the write successful. The system has no way to know whether the failure is transient or permanent. What is not true? I stated an opinion. It is IMHO not the job of the OS to keep data indefinately in RAM just because there might be a chance the storage comes back at some undetermined point in the future. We can fail the write and the application/user can try again from scratch if storage comes back. > That’s a different solution than keeping the dirty pages. Instead you discard the pages and lock to failure. IIRC that’s what openbsd implemented in the wake of fsyncgate. Yes it is different, that's why I said keeping the dirty pages clearly as the paper showed causes all kinds of issues leading to data loss or corruption. I don't see this as a good tradeoff. If OpenBSD went that route then that seems like a much safer and saner decision to me. After re-reading my original comment I see where I wasn't clear enough. Both marking a page as clean and dirty can cause issues if the page is not evicted upon fsync failure. The fact that different filesystems behave differently (some mark clean, some dirty) just makes the problem worse.
- toast0 4y ago
- iforgotpassword 4y ago> Our findings show that although applications use many failure-handling strategies, none are sufficient: fsync failures can cause catastrophic outcomes such as data loss and corruption. That makes it seem like an immediate abort might be the best action in most cases? Handling it wrong and then chugging along might amplify any corruption that has happened. It might obviously depend on the application and use case, but I'd like to think projects like pgsql put a lot of effort into getting this right after fsyncgate. I've read quite a bit about it after that incident, but ultimately decided I'm too stupid to get that right and roll the "log error and bail out" route ever since.
- formerly_proven 4y agoIf you see EIO 99.999 % of the time you want to do just that - stop anything you're doing, don't attempt any further writes. Chances are they'll fail anyway.
- krkoch 4y agoThe storage just bought the farm, EIEIO.
- masklinn 4y ago> ultimately decided I'm too stupid to get that right and roll the "log error and bail out" route ever since. That's exactly where postgres ended up: https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit;h=9ccdd7f66e3324d2b6d3dec282cfa9ff084083f1 https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit... > PANIC on fsync() failure.
- eis 4y agoThe paper explains that even crashing after fsync fails you can end up with lost or corrupted data because after the next start you might get the wrong data from the page cache which survives the crash as it's not part of the programs memory but in the kernel. I recommend reading the paper or watching the video, it is very interesting.
- formerly_proven 4y agoIIRC Linux itself has only been reporting asynchronous writeback errors via fsync for a few short years, meaning before that basically any database that wasn't using O_DIRECT would miss I/O errors under memory pressure (or from out-of-process writebacks in general, e.g. root invoking sync). I looked into this stuff before postgres's fsyncgate, before "how are I/O errors actually handled in Linux, anyhow?" got attention, and walked away with the notion that anything other than O_DIRECT is best-effort-probably-works-most-of-the-time on a good day, and O_DIRECT's semantics are basically an unknowable opaque mixture of what drivers and hardware do and expect. There were some papers looking at error handling within Linux file systems at the time and they found a large number of issues in pretty much all of them. As far as I know, all efforts in the area of durable I/O are still focused on the notion of synchronizing I/O (fsync/fdatasync and equivalent), while many databases don't actually care about that too much and would rather want barriers instead. The kicker is of course that hardware (when honest) actually uses barriers and not block synchronization, and the databases that are journaling filesystems of course also use barriers and not synchronization to implement journaling. It struck me as a distinctly classic API-to-real-world mismatch.
- the8472 4y agoio_uring already has IOSQE_IO_DRAIN which sounds like it's currently implemented as stalling the IO pipeline, but maybe it could be translated to hardware barriers instead in some circumstances (e.g. when the surrounding IO is all O_DIRECT).
- formerly_proven 4y agoThat seems to me like it's not on the right abstraction layer, the drain flag sounds like it's just a barrier for the kernel's threadpool.
- eis 4y agoAfter decades of issues with the storage layer and even some of the most popular programs written by top notch developers having bugs due to the problematic nature of the APIs and filesystems involved I wish a completely new storage API would emerge. Something that exposes an asynchronous (and synchronous build upon it) API with ACID semantics. Filesystems are nothing more than specialized databases but they don't expose the necessary interface to use them as such. We need an API that is dead simple and hard to misuse with clearly defined semantics and guarantees but lets seasoned developers still exploit the hardware to its fullest with additional work. Hope dies last I guess :)
- GordonS 4y agoWindows used to have an almost unknown "transactional file system API" that sounds similar to what you're asking for. I think it was recently deprecated, but I don't know why, or what the history of this API is. Might be interesting to read up on it!
- eis 4y agoYou are probably thinking of WinFS https://en.wikipedia.org/wiki/WinFS https://en.wikipedia.org/wiki/WinFS > In 2013 Bill Gates cited WinFS as his greatest disappointment at Microsoft and that the idea of WinFS was ahead of its time, which will re-emerge
- GordonS 4y agoNope, I was thinking of transactional NTFS, which AFAIK is still supported, but deprecated. I'd totally forgotten about WinFS - I remember it being touted years back, and it sounded great... and then it got cut before the next version of Windows was released. I was pretty disappointed, as at the time it really sounded like the future of file systems. Maybe just ahead of it's time.
- hyc_symas 4y agoYes, deprecated. I thought it had already been dropped but apparently not yet. https://docs.microsoft.com/en-us/windows/win32/fileio/about-transactional-ntfs https://docs.microsoft.com/en-us/windows/win32/fileio/about-...
- chrsig 4y agoOn macOS, most likely not[0]. from the macOS fsync manpage: > fsync() causes all modified data and attributes of fildes to be moved to a permanent storage device. This normally results in all in-core modified copies of buffers for the associated file to be written to a disk. > Note that while fsync() will flush all data from the host to the drive (i.e. the "permanent storage device"), the drive itself may not physically write the data to the platters for quite some time and it may be written in an out-of-order sequence. > Specifically, if the drive loses power or the OS crashes, the application may find that only some or none of their data was written. The disk drive may also re-order the data so that later writes may be present, while earlier writes are not. > This is not a theoretical edge case. This scenario is easily reproduced with real world workloads and drive power failures. > For applications that require tighter guarantees about the integrity of their data, Mac OS X provides the F_FULLFSYNC fcntl. The F_FULLFSYNC fcntl asks the drive to flush all buffered data to permanent storage. Applications, such as databases, that require a strict ordering of > writes should use F_FULLFSYNC to ensure that their data is written in the order they expect. Please see fcntl(2) for more detail. [0] https://twitter.com/marcan42/status/1494213855387734019 https://twitter.com/marcan42/status/1494213855387734019
- xyzzy_plugh 4y agoI said this elsewhere but, in isolation there will always be failure scenarios where recovery is impossible. There are plenty of verification strategies to detect failures, and combined with redundancy, you can reduce the probability of application failure in the face of fsync failures or other similar failures. But you can never eliminate failures. If your storage gives up the ghost, it's game over. Distributed systems are the closest we've gotten to resilient, durable storage. Redundancy, external verification, quorum. Sometimes the distributed system lives in a single box on your desk.
- hyc_symas 4y agoThe description of LMDB's behavior and subsequent analysis are flat wrong. https://twitter.com/hyc_symas/status/1558909442737012736 https://twitter.com/hyc_symas/status/1558909442737012736 To assume that any newbie has hit upon a potential failure condition that we didn't already anticipate and account for in LMDB is frankly laughable.
- simonz05 4y agoThe paper analyzes how file systems and PostgreSQL, LMDB, LevelDB, SQLite, and Redis react to fsync failures. It shows that although applications use many failure-handling strategies, none are sufficient: fsync failures can cause catastrophic outcomes such as data loss and corruption.
- XorNot 4y agoThis sounds a lot like we need to come up with the correct API for this and switch to it.
- GordonS 4y agoI wonder if such an API would require hardware support in order to remain performant?
- eis 4y agoI'd argue some modern hardware e.g. NVME SSDs already expose a hardware interface that would let a new transactional storage API be faster than what we have right now. Unfortunately it is a rare example in a see full of specs that do their best to avoid providing clear semantics and guarantees like atomic sector writes. I guess the next study would need to look into what manufacturers actually fully adhere to the spec. They've not exactly shown best behaviour in that regard in the past (lying about fsync etc) :(
- xyzzy_plugh 4y agoWithout the hardware being open, how do you prove this? In a past life I concluded and argued that there's no point in assuming hardware tells the truth. You have to expect the worst and have a plan for handling it, even if that plan is giving up and failing. Importantly this means you can't claim stronger semantics around e.g. atomicity. You can certainly work around most if not all issues, using redundancy, verification and distribution. But in isolation you cannot even properly observe extreme failure scenarios, you can only reduce their probability, and even that is limited.