3 ms·
There have been bugs involved on both sides. On the kernel side the error reporting was not working reliably in some cases (see [1] for details), so the applic
by pgaddict 8y ago
There have been bugs involved on both sides.
On the kernel side the error reporting was not working reliably in some cases (see [1] for details), so the application using fsync may not actually get the error at all. Hard to handle an error correctly when you don't even get notified about it.
[1] https://www.youtube.com/watch?v=74c19hwY2oE https://www.youtube.com/watch?v=74c19hwY2oE
On the PostgreSQL side, it was the incorrect assumption that fsync retries past writes. It's an understandable mistake, because without the retry it's damn difficult to write an application using fsync correctly. And of course we've found a bunch of other bugs in the ancient error-handling code (which just confirms the common wisdom that error-handling is the least tested part of any code base).
- loeg 8y agoRight; I've been following the headlines on this saga for years :-). > On the PostgreSQL side, it was the incorrect assumption that fsync retries past writes. That's the part I'm claiming is a Linux bug. Marking failed dirty writes as clean is self-induced data loss. > It's an understandable mistake, because without the retry it's damn difficult to write an application using fsync correctly. This is part of why it's a bug. Making it even more difficult for user applications to correctly reason about data integrity is not a great design choice (IMO). > And of course we've found a bunch of other bugs in the ancient error-handling code (which just confirms the common wisdom that error-handling is the least tested part of any code base). +100