8 ms·
Honestly, this doesn't sound like a bad batch of drives or the like -- sounds like they weren't doing scrubbing on their RAID. In case it's helpful, and for ge
by xb95 13y ago
Honestly, this doesn't sound like a bad batch of drives or the like -- sounds like they weren't doing scrubbing on their RAID.
In case it's helpful, and for general knowledge dissemination:
What likely happened is that a drive "failed". This is usually when the RAID card decides that a drive has had enough command errors that it fails the drive. It may actually be fine, and just had a spate of bad responses. You might try to online the drive again and let it rebuild, but that's debatable.
At any rate, so they replaced the drive. That's fine. But then, to rebuild the RAID back to optimal state, it has to read all of the data off of the other drives. Here's where a bad scrubbing policy bites you -- because if those drives have any sectors that have gone bad or problems with the hardware, those drives might fail as soon as the rebuild runs.
Scrubbing should be done regularly (weekly?). What it does is, in essence, test every sector on all of the disks in the array to make sure that all of them are still fully functional so that -- if there is a failure -- you're pretty sure you can rebuild.
The downside of scrubbing is that, for better or worse, it does exercise your disks fairly heavily. Also, if you don't have a suitable trough period then you might even find it difficult to have the available I/O bandwidth to do it.
That said, you should do it if you're not.
- achille2 13y agoNote that they did periodic backups, which would have scrubbed the same sectors as a full resync, so that's not the case.
- xb95 13y agoNot really. A backup would be a filesystem level event, so at best it would only access 1/N of the sectors (where N is the duplication level). It isn't guaranteed to access all of the sectors the data resides in -- and also, it may have hit the kernel's filesystem cache or the on-disk cache or something like that, too, bypassing the physical media entirely. Also keep in mind that a RAID rebuild is a media level event, it doesn't just copy files. It will copy over the "empty" sectors, too, because it just has no idea what's what on the disk. It faithfully recreates the raw bytes on the media. The disk might have failed in a part of the disk you aren't actually using and never noticed because all of your actual data is intact -- the the RAID can't really rebuild without them.
- lmm 13y agoI'd recommend ZFS for working with these kind of arrays; scrub is a command that's easy to regularly schedule, and will only use idle I/O bandwidth (helps that the RAID functionality is integrated with the whole filesystem). If you have an infrequent scrub policy and hit bad sectors on rebuild, it can detect checksum failures and mark specific files as corrupted, rather than declaring all your disks defective. Traditional linux md raid behaviour is particularly bad in this regard: if you have a raid6 configuration and haven't been scrubbing, and then have a single disk failure, all your disks will have a few random isolated bad sectors (i.e. sectors that will URE when you attempt to read them) on, but since you still have one disk's worth of parity it's possible to recover all your data with no downtime (and with ZFS raidz2 this is what would happen). But with md raid as soon as you hit those bad sectors during the rebuild it will consider those drives as failing and kick them out of the array, and since all your drives have at least one bad sector on that means it's impossible to recover the array.
- xb95 13y agoWell, scrubbing with md is pretty easy, too. That said, ZFS sounds great, and I've been meaning to learn it sometime. The big thing holding me off has been just -- well, how much time I've put into learning md and such, and how I would find it hard to justify going back to not knowing much (since this is my livelihood). If you have any tips on "read this, it's a good intro" stuff (slanted to Linux, as that's my cuppa), I'd welcome them.
- lmm 13y agoI'm no expert, just a satisfied user; I got into it from the FreeBSD side, and followed their tutorials, but I don't know how well they'd apply to Linux. In general I like the Gentoo and Arch wikis, particularly for slightly lower-level things like this - even if one is using a different distribution, they tend to take the time and explain the concepts in a way that's mostly distribution-independent. I will say that ZFS felt very coherent and even - dare I say it - easy. I was worried that merging layers that I'm used to thinking of as separate would mean a loss of control, but I was able to put together all the layouts I wanted. The various commands with a single interface felt a lot like using lvm2 - different tools but under a unifying structure with the same parameters - and a lot of what it does is lvm-like. So my best tip is to approach it more as an LVM that can also do raid-like functionality and be mounted directly, rather than as a raid system that includes volume management. Sorry if that's not terribly helpful - as I said, I'm not an expert, but I wanted to at least give you some kind of reply.
- fulafel 13y agoThey didn't mention RAID, and they talk about using SSDs. The failure mode you describe (timeouts) is one typical for spinning rust drives, and not for SSDs.
- xb95 13y ago"we obviously store data in mirrored mode on several servers" They then asked the provider to swap the disk. They had a RAID1; it's a common enough setup. SSDs fail in many of the the same ways as spinning disks, they just add a few more failure cases.
- VLM 13y agoTypical for a killer backplane / killer power supply. Feeding 15 or so volts out the 12 volt line usually has negative results after a couple hours I went thru something like this circa 1996. It was not fun at all. Sometime similar happened with some networking gear at a previous employer. "I know we have pri/sec control cpu boards, but they're burning out as fast as I can swap a new one in!"
- andrewcooke 13y agoeveryone seems to be using this to push their favourite file system, which is fine and all, but if you're using software raid on linux this is the kind of thing you need: #!/bin/bash # # This script checks all RAID devices on the system # http://en.gentoo-wiki.com/wiki/Software_RAID_Install#Data_Scrubbing for raid in /sys/block/md*/md/sync_action; do echo "check" >> ${raid} done (the link referred to seems to be down for me at the moment - i just took this from my main machine). then add a crontab entry to run it once a week or so: 0 0 * * 0 /root/bin/scrub-raid.sh & you can check the script by running it by hand and then doing cat /proc/mdstat where you'll see something like: Personalities : [raid1] [raid0] [raid10] [raid6] [raid5] [raid4] [linear] md1 : active raid1 sdb1[0] sdc1[1] 976760640 blocks super 1.0 [2/2] [UU] [=>...................] check = 6.0% (59419520/976760640) finish=191.3min speed=79902K/sec bitmap: 0/8 pages [0KB], 65536KB chunk finally, my notes on this - http://www.acooke.org/cute/ScrubbingR0.html http://www.acooke.org/cute/ScrubbingR0.html
- anigbrowl 13y agoScrubbing should be done regularly (weekly?) I don't think this applies here. They just moved over to their new RAID setup last Saturday, so they had a major failure within about 72 hours of that transition.
- xb95 13y agoThey use a managed hosting provider, which means that there is no guarantee that the disks are new. When one customer finishes with disks, they get wiped (hopefully) and passed on to the next customer. Even if they are new, though, you never know if it's completely good until you've tested every block on the device. These days most format operations don't actually wipe out the whole disk (takes ages!) so even formatting on a brand new disk won't tell you if it's valid. That said, the odds of having multiple disk failures on a brand new RAID in 72 hours are pretty low. If that's the case, I'd start to suspect a bad batch of disks and/or bad backplane or other hardware (bad power?) that is causing the hardware to fail.