4 ms·
It's interesting how the proper handling of local storage faults is now also recognized to be all the more critical for global replicated systems—that local fau
by _vvhw 5y ago
It's interesting how the proper handling of local storage faults is now also recognized to be all the more critical for global replicated systems—that local faults do propagate across distributed systems.
For example, when not using O_DIRECT, an ack back to the consensus protocol, for local data that was recovered from the log at startup, and not in fact made durable (only marked clean in the kernel page cache after an fsync failure), could cause a quorum swing after the next reboot, and global data loss.