3 ms·
> Every computer component can fail in arbitrary ways, including drives. > > If you’re not robust against that, then when things like fsync fail, then you’ll lo
by pgaddict 7y ago
> Every computer component can fail in arbitrary ways, including drives.
>
> If you’re not robust against that, then when things like fsync fail, then you’ll lose availability and/or data.
The fact is that often the I/O issues are temporary, and those situations are becoming more and more common (think running out of disk space with thin provisioning, networking issues with NFS, etc.). So it might be quite valuable to handle those issues gracefully, without essentially crashing the database (which is pretty much what PANIC does).
> Even though Linux’s fsync behavior is clearly broken, it is far from the craziest behavior I’ve seen from the I/O stack.
"clearly broken" might be bit too harsh, but it certainly makes it way harder to use.
> Anyway, the main lesson here is that untested error handling is worse than no error handling. They should have figured out how to test that this path actually proceeds correctly (on real, intermittently failing hardware) or just panicked the process.
That's true, of course. It's a sad fact of life that error paths are the least tested part of almost any code base. It's however also true that testing I/O errors is pretty difficult to do (especially before "error" dm target), and a significant part of the fsyncgate was that kernel was not reporting errors reliably. It's also true that the behavior is somewhat filesystem-specific (some keep the data in page cache but marked as "not dirty", some will discard the data, ...). That makes testing pretty hard.