4 ms·
Is bitrot something the average hard drive user even needs to worry about? I know the hard drives themselves at the hardware level implement checksums. Is it re
by chadly 13y ago
Is bitrot something the average hard drive user even needs to worry about? I know the hard drives themselves at the hardware level implement checksums. Is it really necessary to also have it at the filesystem level?
I am legitimately asking because I have a good 2TB of family photos on hard drives and spooky stories about random bits flipping freak me out.
- theatrus2 13y agoRunning ZFS for many years has shown, yes, this happens, and will continue to get worse as density goes up.
- olavgg 13y agoI've been running ZFS on several servers with tenfolds of TB's with data. I see checksum errors every month, from a single bit to several megabytes.
- booi 13y agoThat sounds like you have something else wrong. With ECC memory on server hardware, we've seen 0 checksum errors in the last 6 months and I've seen only 2 ever. A typical server has 136TB of raw hdd space and we get about 71TiB usable. It's about 80% full.
- sp332 13y agoDensity is exactly the problem. The chances of an unrecoverable error are around 1 per trillion! But oh wait, you have 2 trillion bits on that disk?
- fiatmoney 13y agoAbsolutely! Over the last ~10 years, with perhaps an average of 6-8 disks at any one time and relatively low intensity use, I've probably had 4 hard drives where the files had checksum errors, before complete failure. But you don't need anything fancy; 2TB is small enough though that you can just buy another HDD and back it up manually.
- lutorm 13y agoExcept that if you don't notice the corruption, you'll just back up the bad data.
- deleted 13y ago[deleted]
- mjb 13y agoYes, both latent sector errors (data that is lost that you don't know is lost) and detectable and undetectable bit errors happen at rates high enough to affect 2TB of family photos. NetApp[1] and the Internet Archive[2] have published good data on this in the past. Scrubbing (periodically reading the data and comparing it to checksums) is one way to get around this. It's very effective against small numbers of sector errors in backups (you should have more than one), and in detecting marginal data on drives. It's less effective against some other kinds of corruption. Another option is to store multiple copies (replication), or additional ECC information that can be used to recover the data if one copy is lost. How much effort you put into this really depends on how much you need or want the data. [1] http://www.cs.wisc.edu/adsl/Publications/latent-sigmetrics07.pdf http://www.cs.wisc.edu/adsl/Publications/latent-sigmetrics07... [2] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.64.1324&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.64....
- wglb 13y agoSuper links, thanks.
- icebraining 13y agoA single bit can be corrected, but there's a threshold per block. As far as I know, the hard drive won't auto-scan blocks that have been sitting still, so it's probably not a bad idea to run a full SMART scan every once in a while.
- Freaky 13y ago> As far as I know, the hard drive won't auto-scan blocks that have been sitting still Seagates at least report their error correction rates via SMART, and make very obvious curves that show them scanning the heads across the disk surface gradually when idle.