4 ms·
This is not the correct modeling approach. All hard drives will fail given enough time, so labeling the failed hard drives as the positive class will bias your
by causalmodels 6y ago
This is not the correct modeling approach. All hard drives will fail given enough time, so labeling the failed hard drives as the positive class will bias your results.
Stuff like this really should be handled using survival analysis.
- mlthoughts2018 6y agoIn fact it screams Bayesian survival analysis, since there is so much prior knowledge both of general hard drive failure rates and SMART stats.
- astrophysician 6y agoI agree and have been doing some similar analysis on the backblaze dataset. I suppose you can use this for prediction but I'm personally just interested in post-hoc analysis and (1) getting better AFR estimates when failure counts are low + (2) exploring time-dependence of hazard functions with different priors (GP priors, etc.). This post and your comment have motivated me to make a post this weekend! Thanks!
- causalmodels 6y agoYou should check out this time-to-event neural network [1]. [1] https://github.com/ragulpr/wtte-rnn https://github.com/ragulpr/wtte-rnn
- jgalentine007 6y agoSMART is a liar sometimes. I have first hand experience with faulty Seagate firmware and Equallogic SANs - where errant statistics caused disks to be ejected from the volumes before you could finish a rebuild. Nothing like watching 40TB of data disappear on multiple installations over the course of a few weeks!
- R0b0t1 6y agoSMART seems to be extremely useless in practice. Manufacturers don't seem to expose actual failure statistics through it, likely for fear of making their product look bad.