6 ms·
I worked at a storage company and the scariest thing I learned is that your data can be corrupt even though the drive itself says that the data was written corr
by steven2012 11y ago
I worked at a storage company and the scariest thing I learned is that your data can be corrupt even though the drive itself says that the data was written correctly. The only way to really be sure is to check your files after writing them that they match. Now whenever I do a backup, I always go through them one more time and do a byte-by-byte comparison before being assured that it's okay.
- jwr 11y agoThis is true. Which is why we really, really need checksummed filesystems. I am very worried that this hasn't made its way into mainstream computing yet, especially given the growing drive sizes and massive CPU speed increases.
- sneak 11y agoFortunately, zfs on Linux is excellent, and is a two-liner on modern Ububtu LTS. (add PPA, install zfs.)
- tshtf 11y agoIs it? I've heard several complains about bugs in FUSE.
- mitchty 11y agoI've a friend that uses it. Can't say it is not buggy, he hit the bug where unlinked files weren't removed. Got to 95% use before finding out he had to reboot/unmount to clean things up. But that means his pool performance is now shit.
- gnoway 11y agoZFS on Linux[0] doesn't use FUSE. [0] http://zfsonlinux.org/ http://zfsonlinux.org/
- tshtf 11y agoThanks, I was unaware. Apparently there is both native ZFS, and FUSE-backed ZFS for Linux: https://en.wikipedia.org/wiki/ZFS#Linux https://en.wikipedia.org/wiki/ZFS#Linux
- barrkel 11y agoI run a 10x3TB ZFS raidz2 array at home. I've seen 18 checksum errors at the device level in the last year - these are corruption from the device that ZFS detected with a checksum, and was able to correct using redundancy. If you're not checksumming at some level in your system, you should be outsourcing your storage to someone else; consumer level hardware with commodity file systems isn't good enough.
- mitchty 11y agoAs a counterpoint I have 6x3TB zfs raidz2 on freebsd at home. I resilver every month and only had one checksum error that turned out to be a cable going bad given it hasn't repeated. Still agree with we need checksumming filesystems though. That and gcc ram to make the data written more trustworthy.
- PhantomGremlin 11y agohad one checksum error that turned out to be a cable going bad given it hasn't repeated I wouldn't assume that it was a cable error. The SATA interface has a CRC check. So the odds are quite high that a single error would simply result in a retransmission. Of course, a plethora of detected SATA CRC errors and resulting retransmissions means that an undetected error could readily slip thru. There should be error logs reporting on the occurrence of retransmission, but I'm not enough of a software person to know how possible / easy it is to get that information from a drive or operating system. OTOH, as you mention later in your post, a single bit error in non-ECC RAM could easily result in a single block having a checksum error. Exactly what you saw!
- wtallis 11y agoHard drives also have spare sectors, so if a defect is detected at one spot in the disk, it will probably never touch that spot again. Simply observing that an error only occurred once does almost nothing to narrow down the possible causes. You have to also be keeping track of all the error reporting facilities (SMART, PCIe AER, ECC RAM, etc.).
- 11y ago
- crayg33k 11y agoThis is why end-to-end data integrity with something like T10-PI is a necessity. The kernel block-layer already generates and validates the integrity for us, if the underlying drive supports it, but all major filesystems really need to start supporting it as well.
- cmurf 11y agoI don't think that's a necessity for all workflows. Just think about it, that would require all of us buying enterprise 520 or 528 byte sector drives to store the extra checksum information, and a whole new API up to the application level to confirm, point to point, that the data in the app is the data on the drive on writes, and the data on the drive is the data in the app on reads. It's not like T10/PI comes for free just by doing any one thing, it implies changes throughout the chain.
- avar 11y ago> The only way to really be sure is to check your files after > writing them that they match. This is assuming that the underlying block device would forcibly flush those queued writes to disk and then re-read them again rather than just serve them up directly from the pending write queue directly without flushing them first. You generally can't make that assumption about a black box, so reading back your writes guarantees nothing. Unless you're intimately familiar with your underlying block device you really can't guarantee anything about writes going to physical hardware. All you can do is read its documentation and hope for the best. If you need a general hack to that's pretty much guaranteed to flush your writes to a physical disk it would be something like: After your write, append X random bytes to a file where X is greater than your block device's advertised internal memory, then call fsync(). Even then you have no guarantees that those writes wouldn't be flushed to the medium while leaving the writes you care about in the block device's internal memory.