4 ms·
Interesting, thanks for posting! Could you talk quickly about why it's interesting to predict drive failure? Is it to understand how many replacement drives you
by Synroc 10y ago
Interesting, thanks for posting! Could you talk quickly about why it's interesting to predict drive failure? Is it to understand how many replacement drives you might need to order in the short term, or is there value beyond stock management of drives?
- matt4077 10y agoNot OP, but: Perfect predictability would obviously be beneficial, in that you could get by without any redundancy. But even imperfect predicability can help you reduce the required number of drives for a set level of security.
- honkhonkpants 10y agoYeah but no. The problem is that absence of smart counters does in no way indicate good health of the device.
- matt4077 10y agoIf the presence of smart counters indicates imminent failure, their absence by definition indicates health. Not perfectly, obviously, but this is about probabilities. Say you want some level of data security (i. e. 99.9% over one year). The formula for the risk of data loss is (r^n) where r is the failure rate of drives and n is the number of drives (assuming independence and using just mirroring). But – if you have a test that can predict failure with some probability s the formula becomes (n*(1-s))^r. which is strictly smaller for any 0 < s <= 1, meaning you get a higher level of expected security (or possibly the same with fewer drives).
- honkhonkpants 10y agoThis is a discrete problem. You cannot have 2.9 replicas instead of 3. The fact that 24% of these drives failed without smart indicators means no test exists which can lower your replication from 3 to 2.
- nitrogen 10y agoCan having ~2.9 replicas be approximated by using an error-correcting code across a large enough number of drives? Isn't that how RAID 5/6/Z work?
- honkhonkpants 10y agoSort of but again no. He problem is if you have really wide coded stripes your I/O costs to reconstruct after a failure will be astronomical, and the probability of one failure increases with the number of participating devices. Besides you could not reshape the stripe in response to anticipated failures without reading and writing the whole thing, in which case you'd be better off in terms of I/O costs just evacuating the device in question.
- matt4077 10y agoWell, by that measure it doesn't even matter if a drive fails with 10% or 5% probability beyond the replacement costs. Because a test with 50% sensitivity effectively halves the failure rate you have to use. Since it's all probabilities, triple redundancy does not guarantee complete absence of data loss. On the other side of the spectrum, single replication might offer a better cost/safety ratio for some applications. A failure that is known in advance is equal to no failure in these term and even when discrete, it will make the difference once in a while.
- kryptiskt 10y agoYou can reduce your risk by requiring that no more than one of the three replicas is on an iffy drive, creating new replicas on healthy drives if that's not true.
- greglindahl 10y agoIn a non-RAID context, for example NoSQL databases that keep 3 copies of chunks of data, knowing about a failure in advance means that you can abandon using that drive slowly, without it becoming an emergency.
- honkhonkpants 10y agoIf you have three replicas, who cares if one fails? Just wait for it to fail and rereplicate from the survivors.
- greglindahl 10y agoBecause there's a risk that another replica will fail. And because you can do the copying more slowly if the failure is predicted and not super-immediate.
- andy4blaze 10y agoBesides stock management, they help use determine the overall health of a Storage Pod or Vault. They also help find trouble with other components. For example, if a backplane or cable were failing, the drives via their SMART stats may notice first. So they SMART stats are part of what we use to evaluate the whole system health.