22 ms·
If it was half failed we would DEFINITELY pull the drive out because it is already 1 drive down out of 20 for half the files. A lot of times the IT guys will m
by brianwski 8y ago
If it was half failed we would DEFINITELY pull the drive out because it is already 1 drive down out of 20 for half the files. A lot of times the IT guys will make a judgement call that a drive is acting funny or slightly off so they just "fail it on purpose" which means yank it and replace with a new drive. We have done this just because a drive is "slow" (slow can mean the drive is having trouble writing data reliably on one attempt), or because some SMART stat looks wonky.
To provide more color, if a 20 drive "tome" (as we call it) is 1 drive down, we don't even wake people up in the middle of the night, but Backblaze datacenter employees replace it when they arrive at the datacenter the next day at 8am. All drives having problems are replaced by 5pm when the employees go home. This is completely business as usual, about 5 - 10 drives fail every day.
However, if 2 drives fail out of 20 (or 1.5 in our example above), pagers go off, people wake up and get out of bed at 3am and start driving towards the datacenter. Or we employ "remote hands" to swap the drives immediately, it depends on the capabilities of the night crew in the datacenter which varies by datacenter. "remote hands" is a contract service where semi-skilled technicians work for the datacenter and we can pay them $80/incident or there abouts to do things you can only do "in person" like replace drives. All the pods (where data is stored) have "base board management" which means as long as they are powered up and online we can log in remotely from home or office to figure out what is going on and fix a variety of problems. AUTOMATICALLY if 2 drive fail we stop sending any data into that "tome" of 20 drives. We have found that writing to drives causes more failures, so not writing to them is safer.
If 3 drives fail, it is instantly a "Red Alert" at Backblaze and a whole lot of official procedures kick in. An "incident manager" is assigned and the whole company's number one concern is to drop EVERYTHING and never sleep again until the Red Alert is lowered to Yellow. We light up a "situation room" (in Slack - our internal chat tool) and information and status is relayed through that.
SIDE NOTE: Backblaze has a relationship with an excellent company named "DriveSavers" who can recover SOME data off of failed drives. This is very expensive (thousands of dollars per drive) so we only do it to test the procedure and then in extreme situations. Three drives down is an extreme situation and extremely rare, so ALL OF THE THREE FAILED DRIVES would be immediately hand carried to DriveSavers even while we rebuild the customer data from parity. Notice Backblaze STILL has a complete copy of the customer data on 17 drives -> But if a 4th drive dies, the hope is we can recover at least one of the drives via DriveSavers thus saving the customer data. (We need at least 17 out of 20 drives in a "tome" to reconstruct the data.) In our experiments, DriveSavers seems to recover about half the drives, or in some situations half the data from a drive (imagine if 1 platter on a drive has a head crash and is destroyed, but the other platters are fine). We have made the decision that it is less expensive (for the same durability) to pay DriveSavers the thousands of dollars rarely instead of increasing parity to allow reconstructing data from 16 out of 20 drives instead of the current 17 out of 20 drives.
- zepearl 8y agoThanks a lot - absolutely interesting/fascinating. You guys should think about writing a short eBook about e.g. general recommendations about setups/analysis/projections & stories about past failures/chain-of-events/etc - I might buy it :)