4 ms·
To be fair, they didn't really understand how ZFS works and failed to set up bitrot detection. A "clean" setup would include those, as well as either a messagi
by 7steps2much 5y ago
To be fair, they didn't really understand how ZFS works and failed to set up bitrot detection.
A "clean" setup would include those, as well as either a messaging system or a regular checkup on how your FS is doing.
- liuliu 5y agoNot only that. A vdev with 15 disks each around 16 (or 18?) TiB with RAID-Z2 just asking for troubles (https://www.zdnet.com/article/why-raid-6-stops-working-in-2019/ https://www.zdnet.com/article/why-raid-6-stops-working-in-20...). Even with 1% disk failures annually (based on Backblaze, for better disks), we are looking at least two disks fail at the same time with 15 disks around 1% chance. With 10 vdevs, that is 10% chance of any of these have failures. They should really move to RAID-Z3 or 8 disks group. With RAID-Z3, we are looking at at least 3 disk fail at the same time with 15 disks around 0.04% chance. With 8 disks, we are looking at 0.2% chance.
- kalleboo 5y agoI've read this article and others like it, and what I don't get is if the URE rates are really that high, why I have I never seen an URE in the bi-weekly scrubs of my pool (184 TB of raw disk), aside from when a single disk was literally going bad?
- sliken 5y agoThere's many contributing factors. Power supplies degrade over time, ability to maintain the correct voltages can be impaired for worst case scenarios, like running all disks flat out. Drives (and even more so older drives) can be vibration sensitive, especially in worse case scenarios involving all drives running flat out. With modern manufacturing tolerances being so low, whatever triggered the first disk in your pool to die is likely to get a second in a fairly small window. So sure, the failures are not equally distributed and you've done well with your 184TB pool. Are you tracking the device errors, or only those that are visible to the OS? But multiple disk failures are not particularly uncommon. My experience in this space was two 16 disk servers that I set up 6 5 disk RAID5s (with one global spare per server). Within one month I had 11 of 32 disks die, and barely managed not to lose any user files, and this was not during the 1st month in production. Scary. I've since moved to pairs of servers cross connected to pairs of 60 disk chassis (16 x 12 gbit connections per chassis) with ten 11-disk RAIDz3, 10 global spares, and 6x3.2TB of NVMe cache per server.
- kalleboo 5y agoOh there are all kinds of reasons drives can cause errors, and you have the bathtub curve. So there's lots to take into account when designing your pool. But the article is using the spec sheet URE rate which I'd assume looks only at the drive and doesn't take into account problems with the computer around the drive or the EOL time after the drive warranty has expired, I'd assume it was the "baseline" error rate. > Are you tracking the device errors, or only those that are visible to the OS? If we're talking URE like the article, that's data-loss on a disk, and the OS would always figure it out, since it would cause a ZFS checksum failure on scrub. In this case it's not my data and not my money, so my preference is 6-drive RAIDZ2 vdevs. We've only had one disk with errors (and that one was migrated from a PC where Windows never reported any errors... of course...). The oldest 2 disks (3.5 years power-on time) have single-digit reallocated sectors in SMART so those are on course to be replaced. I'm just curious since the argument in the article doesn't add up in my eyes. > Within one month I had 11 of 32 disks die, and barely managed not to lose any user files, and this was not during the 1st month in production Wow, that is some terrible luck!
- mustache_kimono 5y agoThey even admit this in the video -- "No one is to blame here except us" which means "Holy shit! We really messed this up." Yes, they should have been scrubbing their pools, but I don't think this was bit rot. Millions of data errors is not what bit rot looks like. This looks exactly like bad hardware.
- ksec 5y agoServetheHome actually reported the hard drive LTT are using had higher failure rate than others. Even from a consumer hardware perspective. And they were suppose to be enterprise drive. It is just bad hardware, bad setup and bad everything all cramped together.
- mustache_kimono 5y agoI've never seen a HDD fail like that, but I'm willing to believe it's possible. I have seen read/write errors because of a buggy implementation of ALPM. But, agreed, whatever it was, if they were paying attention, they could have diagnosed and remediated well before they had any data loss.