5 ms·
We never really worried about write endurance with dynamic RAM, but it would seem to be a lot busier than disk I/O at least in some server applications.
by JudasGoat 8y ago
We never really worried about write endurance with dynamic RAM, but it would seem to be a lot busier than disk I/O at least in some server applications.
- marshray 8y agoWe need reviewers report a new metric: time to failure at continuous max write throughput. For some SSDs, it's under a week.
- Moral_ 8y agoThe write to failure is usually reported in the drive specs. https://www.intel.com/content/www/us/en/products/memory-storage/solid-state-drives/data-center-ssds/optane-dc-p4800x-series/p4800x-750gb-aic.html https://www.intel.com/content/www/us/en/products/memory-stor... See Endurance Rating (Lifetime Writes)
- marshray 8y ago"Endurance Rating (Lifetime Writes) 41.0 PBW" "Sequential Write (up to) 2200 MB/s" 41e15 / 2.2e9 /60/60/24 = 215 days sequential write to failure "Mean Time Between Failures (MTBF) 2 million hours" 2e6 / 24/365.25 = 228 years MTBF So it seems the MTBF is being stated at 0.26% average write utilization. [corrected math]
- mjevans 8y agoFor some applications that might be worth it.
- deleted 8y ago[deleted]
- Retric 8y ago0.26% * 2200 MB/s = 5.72 MB/s for 24/7/365 which seems about right for most users. A 750 GB drive assuming you want to store the data for 48 hours can only write (750 GB/24/60/60*1000) = 4.34 MB/s on average. Dropping that to even 1 hour still gives reasonable lifetimes.
- dmayle 8y agoThat's not how MTBF works. It's not how long you expect the drive to last. The number is how many failures on average you get for the number of service hours in use. It's really only useful in aggregate. For a MTBF of 2 million hours, that means that on average, if you have one thousand drives, then you should expect one drive failure every 2000 hours, or 83 days (1k * 2k hours = 2M hours)
- namibj 8y agoMTBF is what you need to take a look at for a mechanical system, if you are concerned about service intervals (in a HA cluster style setup, where you can handle fixing after break down, without downtime), and for e.g. continuous operation of optical disc and magnetic tape drives, as those wear out over time (though even laptop BD-R drives reach 1 year active spindle MTBF). There is determines how often you have to go to the system and swap faulty drives, to have at least x% working, and over how much time you can budget the CAPEX. Of course this breaks down at an MTBF of over 50 years, as thosre rarely mention exotic failure modes, and don't actually have an MTBF of 50+ years over the life, but an annualized failure rate corresponding to 50+ years MTBF, measured over the first couple of years or even the first year. For non-wet-electrolytic-capacitor-using computing, one can calculate a temperature-dependent MTBF in the 5-25 years range, mostly depending on how bad the chips are hit by electromigration and similar aging in the semiconductors. This is incidentally a reason why I miss clock speeds for different processors as reported by overclockers to at least in some cases extrapolate the life due to electromigration, as there is a formula with like iirc 2 parameters, which gives a temperature (and maybe voltage) dependent lifetime/MTBF for this semiconductor device. I'd likely sttrive for about 5 years MTBF on the processor, if speed is of concern and reliability/uptime not in the foreground.
- marshray 8y ago> For a MTBF of 2 million hours, that means that on average, if you have one thousand drives, then you should expect one drive failure every 2000 hours That assumes your drives fail with a constant independent probability, like nuclear decay events (Poisson distribution). The reality is more like https://en.wikipedia.org/wiki/Bathtub_curve https://en.wikipedia.org/wiki/Bathtub_curve . MTBF is not a good metric for complex systems used under wildly-varying load conditions, but ... it's a metric.
- surajrmal 8y agoYou're not taking into consideration write amplification. Assuming 3x write amplification, the actual userspace writes would be 1/3rd.
- pilsetnieks 8y agoAre you sure? A quick back of a napkin calculation seems to suggest that a fully saturated SATA3 connection would still be a measly 36 terabytes in a week. It seems quite low for any practical purpose. I don't doubt that there probably are some tiny shitty drives that will conk out after a week like that but are there any reasonably popular drives like that?
- djsumdog 8y agoWhat about nvme/pcie drives? They can push considerably more data.
- marshray 8y agohttps://twitter.com/marshray/status/870087265984233476 https://twitter.com/marshray/status/870087265984233476 This was a drive by a top brand NAND manufacturer.
- opencl 8y agoThe 500GB 970 Pro is rated for 2300 MB/s sequential writes and 600 TB write endurance. That's about three days to exhaust the write endurance. Latest high end SSD model from the leading manufacturer. Not that it could actually come anywhere near sustaining that throughput for three days straight. https://www.anandtech.com/show/12674/samsung-announces-970-pro-and-970-evo-nvme-ssds https://www.anandtech.com/show/12674/samsung-announces-970-p...
- AlphaSite 8y agoThe perf isn’t really that good for ext extended periods.