8 ms·
SSD Failures in Datacenters
- chmaynard 10y agoBad URL?
- khc 10y ago"An error occurred while processing your request. Reference #50.d66d717.1467850467.45a102d" I supposed their ssd failed
- sua_mae 10y agoAkamai failed
- sharva 10y agohttp://dl.acm.org/citation.cfm?doid=2928275.2928278 http://dl.acm.org/citation.cfm?doid=2928275.2928278
- RUBwkVjwLsDKgPw 10y agoAll or most of the ssds used in this study are used as a cache in front of hdds, so take results with a grain of salt.
- PostThisTooFast 10y agoSSDs HDDs
- kyledrake 10y agoSummary?
- bluecmd 10y agoLike any paper, this too has a summary: 6. Concluding Remarks This paper presents an extensive characterization of SSD failures using field data.We first show that SSD failure rates in the field can be very different from what vendors specify. Next, we identify and quantify four types of SMART failure symptoms exhibited by the SSDs, and provide characteristics of symptom occurrence, intensity and progression rate. We show that despite their presence, the symptoms alone cannot be a sufficient indicator of failures. We have also studied the impact of multiple provisioning and operational factors across different layers of the datacenter hierarchy on SSD reliability. Many of these factors are individually influencing, and can also interact with each other in complicated ways, to impact SSD failures. We have used machine learning and graphical model based approaches to systematically consider the impact of multiple influential factors towards answering the what, when and why of SSD failures. We believe the insights gained from this paper can greatly influence the design, provisioning and operational decisions for SSDs in datacenters.
- SFJulie 10y agoCompared to HDD SMART yields quite a lot of false negative/positive hence it is proving harder than HDD means to detect failures. Failure rates seems to not fit the models and be a tad higher than expected. Most SSD are reliable, the only problem is they are failing in a way that is tough to detect leading to consequences we may ignore. (falsely positive functioning SSD in operations and we may experience silent corruptions) Given the "too many knobs" of the SSD they can defect in non predictable snowballing effects (chaotic and dramatic). Basically this study confirms that IEEE standards for rating MTBF are to be reconsidered drastically and that our models of failure for SSD are far from totally being well understood and that SSD is in production while all costs are not totally yet known. The Cost of SSD specific failures is not yet known and they urge SSD makers to begin studying the effects of their "knobs" on reliability Even though it is written in a very scientific neutral tone trying not to scare people, it basically can be used as a strong evidence to ban SSD from critical systems. It basically is an heavy blow to SSD industry since it attacks it on the costs model that is based on its expected reliability and says that these drives are basically still an unknown territory when it comes to failures.
- RP_Joe 10y agoWhats the point in posting paywall links?
- OriginalPenguin 10y agoIt's not a paywall. To read the article click the link: "Full Text: PDF".
- gondo 10y agoWhy not linking to PDF directly then?
- paraxisi 10y agoIt's less opaque - there's an abstract, references/citations, etc. Also, I suspect I'm not the only one who doesn't want to be directly linked to things like PDFs that have a history of causing security issues.
- Naga 10y agoI'm on mobile now so I appreciate being able to read the abstract without loading a whole PDF.
- PuffinBlue 10y agoAs long as you're not appreciating that because you're thinking the page with the abstract would be smaller, the PDF is about 30% the size of the web page.
- rbanffy 10y agoEven if it were, this specific paywall is well worth it.
- Jaruzel 10y agoFrom the HN FAQ: Are paywalls ok? It's ok to post stories from sites with paywalls that have workarounds. In comments, it's ok to ask how to read an article and to help other users do so. But please don't post complaints about paywalls. Those are off topic. -- This will be why you are being downvoted.
- pcunite 10y agoCan someone summarize the results? Should I continue using my SSD, or no? :-)
- Dylan16807 10y agoA typical annual failure rate near but under 1%, so use it like any drive by having regular backups and expecting to have to use them at some point.
- kyrra 10y agoFor those interested, Google published a similar report a few months back. https://www.usenix.org/conference/fast16/technical-sessions/presentation/schroeder https://www.usenix.org/conference/fast16/technical-sessions/... https://news.ycombinator.com/item?id=11188445 https://news.ycombinator.com/item?id=11188445
- spudlyo 10y agoI worked at a company that migrated to SSDs on hundreds of HP ProLiant servers. The performance was great, it really improved our DB latency. Unfortunately the RAID 1+0 setup that was carried over from the spinning rust setup didn't work for our most common failure case; corrupted writes. The HP RAID controller failed to detect corrupted writes on what I seem to remember were Intel SSDs. The only way we learned about failed SSD drives is when we started catching large numbers of DB page checksum errors, and by that time it was too late to do anything about it. Operationally it sucked, swapping out drives is a lot easier maintenance than having to rebuild a DB from backup.
- Veratyr 10y agoIs there a reason you were using hardware RAID rather than a software based RAID with checksumming like ZFS? (I mean this as an honest question, I'm interested to know if there are performance or reliability gains to be had from a hardware controller)
- spudlyo 10y agoThe oft repeated performance reason is that hardware RAID controllers have battery backed write cache, which really improves durable write performance; something traditional relational DBs do quite a lot of. Also, these machines were built by DB people, who have traditionally placed a lot of value in hardware RAID controllers, and to be fair, the HP controllers worked well with HDDs. Also ZFS wasn't a realistic option for production Linux in 2013. Me personally, I'm not a fan of exotic proprietary hardware, I much prefer working with JBOD disk setups and software RAID where appropriate.
- creshal 10y ago> The oft repeated performance reason is that hardware RAID controllers have battery backed write cache, which really improves durable write performance; something traditional relational DBs do quite a lot of. Don't all enterprise SSDs come with their own backup capacitors?
- mikevm 10y agoAnother interesting paper on the issue of using parity RAID with SSDs: http://pages.cs.wisc.edu/~kadav/new/pdfs/diffraid-hs09.pdf http://pages.cs.wisc.edu/~kadav/new/pdfs/diffraid-hs09.pdf
- more_corn 10y agoI did a large production SSD deployment a few years ago on a db fleet backing a large video hosting site (probably the first one you think of). tl:dr the benefits were astounding. The failure rates were super low. We saved something like 2 million dollars by significantly increasing the I/O performance of the fleet preventing the need to re-shard. Choosing a drive We went with an Intel MLC (consumer class) drives. No other drive had such low DOA rates and good match between performance and price. We set the max lba to 80% of available capacity. (Actually recommended by Intel) This change eliminated the pathological case where a full drive continually overwrites the same sectors. The consumer drives suddenly had a lifespan comparable to enterprise drives (though a slightly slower read speed- which was fine because we needed balanced reads and writes) We also exported and monitored the wear leveling stats (available via a SMART value) so we won't run into the case where they all unexpectedly wear down at the same time. The projected lifespan looked really good and in most cases exceeded that of the servers. (in retrospect this turned out to be true) Raid config We ran in Raid 0 (because I love to live on the edge- No, actually because we split the boot, data and bin log volumes and data integrity was preserved by redundancy within the the fleet.) The conversation I had with the dbas about this was one of the most eye opening conversations of my career. Turns out if a replica failed or lagged too much they simply tossed the data and started a new one. They didn't actually need or want RAID10 and that changes the economics of the picture significantly. This would have been 4 years ago now. We deployed several thousand drives. Of that only one ended up with abnormally high wear level. (I proactively threw it out when we decommissioned the hardware to prevent a problem for the next guy). I had a handful of drives which didn't pass the pre-service checks and a another handful that failed in production. Compare that to 10k spinners which appear to have an annual failure rate of 1 in 20. SSD Failure profile: We would typically see a drive failing to respond properly (announcing a crazy size, not showing up at all, failing to save its max LBA). Most of this happened in the pre-production check phase. Post deployment failures you can count on one hand (I have the normal number of fingers ;-) Qual process We qualified 6-8 different manufacturers, high end enterprise to the dirt cheap vendor with a three letter name now synonymous with garbage (I won't even say the name here). The cheap drive had high DOA rates, and the unsettling tendency to spontaneously reset itself causing a several second pause (not so good for our uses, but might be tolerable in other circumstances)
- smartbit 10y agoslides http://www.systor.org/2016/slides/ssd_failures_systor2016.pdf http://www.systor.org/2016/slides/ssd_failures_systor2016.pd...