4 ms·
I read that heat has yet to be proven to detriment a drive's life or performance within reason. Sure, 5000 degrees would be bad, but there apparently is no evi
by biturd 15y ago
I read that heat has yet to be proven to detriment a drive's life or performance within reason. Sure, 5000 degrees would be bad, but there apparently is no evidence that cooling drives matters.
I can't source the article, but seem to recall is was google related with the data coming from their experience with drives in their data centers. It possibly could have come from BackBlaze, though that is a distant second.
I used to run a small ISP, and found that I had a 10% failure rate on average of anything, before it was even used. Order 100 new servers, 10 will be dead on arrival. Order 1000 drives, 100 will be DOA. 100 network cables, all will work, that was about the only deviation.
With drives, the failure rates were better, however, I never allowed drives to be put into use that were not internally "certified". I would use SoftRAID, which is a Macintosh application, though I am sure there are equivalent on Linux and Windows. While this is nothing more than software RAID for the Mac, many of my servers were not Macintosh. I didn't use the RAID capacity of the software, rather using the drive certification feature instead. It takes about 8 hours per 1TB drive to "certify" it. This will run through every sector on the drive, and make sure it is ok. The softwares ability to "predict" that a drive was about to fail, or susceptible to failure, was very good. I would never use a drive that didn't pass.
One handy feature which I did run on non servers with the software, is that it can monitor the drive and tell me that is suspects failure to be imminent. This is not SMART, though SMART was a first line test, the software went deeper and logged aspects of the drive over time.
Once I started doing this, drive failure went to near nothing. I can't really measure failure in percentage or time, as we would outgrow the capacity of the drive, or the server would be taken out of commission for a faster server before a failure ever happened. Those drives were then given to friends or sold. As a result, we never had an in use failure once we started using this method of testing. 42U worth of server space, with a mix of 1U, 2U, 4U, and a few 6U NAS's, plus switching gear.
Do keep in mind, we were a small local ISP, and could take the time to perform these tests. Someone like BackBlaze has their system created in a way to deal with failure as part of the operation, and it would be a waste of time to perform these tests.
- deleted 15y ago[deleted]
- ComputerGuru 15y agoThanks for sharing your story, and an upvote for all the details. But don't you think it's a little bit dishonest to sell perfectly healthy drives that you knew to be on the verge of failure?
- Kliment 15y agoThe way I read it, they sold the healthy drives that they had outgrown and needed to replace for operational rather than technical reasons.
- joshu 15y agoIf you shopped at Fry's you'd have 10% bad network cables, too.
- rhplus 15y agoThe Google paper you're referring to is here[1]. Their results show that very high temperatures do affect older drives (> 40ºC, 3+ years), but temperature variation within lower and mid ranges shows little correlation with failure rates. The important result from Google's and others'[2] research is that lifespan for particular drive models is essentially bimodal, with the first peak during the first couple of months - strongly correlated with manufacturing batch, not just make/model - and the second peak after 3+ years due to age related causes. It's possible that people who experience entire arrays of drives failing within the first few months of use have chosen a particular batch of drives that suffer from high 'infant mortality'. One recommendation is to build arrays from drives with different manufacturing batch number or - if the RAID implementation allows it - use completely different makes and models. [1] http://research.google.com/archive/disk_failures.pdf http://research.google.com/archive/disk_failures.pdf [2] http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=1285441 http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=1285...