3 ms·
More than half the time in my experience it’s usually a failing sata cable or power-supply if the hard drive continues to functions on boot but eventually goes
by redflame7 6y ago
More than half the time in my experience it’s usually a failing sata cable or power-supply if the hard drive continues to functions on boot but eventually goes to degraded due to checksum/read errors. Sometimes the data is just fine but a read error occurs between the drive and controller which ZFS interprets as a failng disk
- 867-5309 6y agomodern drives have onboard diagnostics which can report errors directly to the OS. I'd be very surprised if a modern file system was unaware of this. I'd also be surprised if users of alternative file systems didn't test a suspected failing drive before concluding it was faulty
- redflame7 6y agoThe error is occurring between the drive and the sata controller on the mobo. There’s also the whole issue of whether the OS or the drive is responsible for error handling. If the the drives are SAS drives, typically the OS uses fire and forget methodologies for write commands. Other drives might not support that and require zfs to handle errors at the software/kernel level
- 867-5309 6y agoreplacing the cable / testing the disk on an isolated system would confirm this
- blueflame7 6y agoIt would but I encountered a situation once, where when I would remove the drives of the array and test the disk 1 by 1 they would all work. It wasn't until all the drives were powered on and under max write stress would the power-supply under-volt to drives randomly making it appear as multiple drives were all failing.
- blibble 6y agothe great thing about zfs is you don't have to care, at all as long as it writes and reads back (most of the time): ZFS will deal with it
- ianhowson 6y agoHard drives lie. SMART data is not thorough and it is not trustworthy. Across dozens of drive failures over the years, I've never seen SMART predict a failure before ZFS detected data loss or the OS reported 'drive is slow'. SMART is not worth the effort to monitor. Filesystems don't read SMART data. You might have a separate daemon which monitors SMART. ZFS checksumming is amazing for this. You know, without doubt, which file(s) are bad. You can still use failing or unknown quality drives because the checksumming will protect you from silent data corruption.
- toast0 6y agoSMART marking drives as failed is usually super late. But if you monitor the raw values for the important parameters (unreadable, reallocated, etc sectors), I've had good luck with replacing drives before software notices. At least for spinning drives; SSD failues were way more rare, but resulted in the drives completely disappearing from the bus, and no reliable prefail indicators. I did have one SSD go through a big reallocation that tanked throughput and the alerts from throughput and the alerts from SMART thresholds fired simultaneously. Of course, that's great for a server farm; in home use monitoring is a lot less structured.
- pixl97 6y agoExpensive enterprise drives are generally better at smart stats. Of course your mileage varies by vendor. Desktop drives on the other hand are completely untrustworthy.