4 ms·
We certainly haven't ruled out Postgres or kernel bugs here. > If you ran into a problem like this on ZFS, for example, you'd have very high confidence about w
by richvdh 1y ago
We certainly haven't ruled out Postgres or kernel bugs here.
> If you ran into a problem like this on ZFS, for example, you'd have very high confidence about whether the disk was at fault
Would we, though? I'll admit to not being that familiar with ZFS's internals, but I'd be a bit surprised if its checksums can detect lost writes. More generally, I'm not entirely sure how practical it would be to add verification at all layers of the stack, as you seem to be suggesting.
We'd certainly be open to considering ZFS in future if it can help track down this sort of problem.
- dap 1y ago>> If you ran into a problem like this on ZFS, for example, you'd have very high confidence about whether the disk was at fault > Would we, though? I'll admit to not being that familiar with ZFS's internals, but I'd be a bit surprised if its checksums can detect lost writes. Yup. Quoting https://en.wikipedia.org/wiki/ZFS#Data_integrity https://en.wikipedia.org/wiki/ZFS#Data_integrity: > One major feature that distinguishes ZFS from other file systems is that it is designed with a focus on data integrity by protecting the user's data on disk against silent data corruption caused by data degradation, power surges (voltage spikes), bugs in disk firmware, phantom writes (the previous write did not make it to disk), misdirected reads/writes (the disk accesses the wrong block), DMA parity errors between the array and server memory or from the driver (since the checksum validates data inside the array), driver errors (data winds up in the wrong buffer inside the kernel), accidental overwrites (such as swapping to a live file system), etc. (end of quote) It does this by maintaining the data checksums within the tree that makes up the filesystem's structure. Any time nodes in the tree refer to data that's on disk, they also include the expected checksum of that data. That's recursive up to the root of the tree. So any time it needs data from disk, it knows when that data is not correct. --- > More generally, I'm not entirely sure how practical it would be to add verification at all layers of the stack, as you seem to be suggesting. Yeah, it definitely could be a lot of work. (Dealing with data corruption in production once it's happened is also a lot of work!) It depends on the application and how much awareness of this problem was baked into its design. An RDBMS being built today could easily include a checksum mechanism like ZFS and it looks like CockroachDB does include something like this (just as an example). Adding this to PostgreSQL today could be a huge undertaking, for all I know. I've seen plenty of other applications (much simpler than PostgreSQL, though some still fairly complex) that store bits of data on disk and do include checksums to detect corruption.
- dap 1y agoSorry I didn’t say it sooner: thanks for sharing this post! And for your work on Matrix. (Sorry my initial post focused on the negative. This kind of thing brings up a lot of scar tissue for me but that’s not on you.)