4 ms·
> Guessing shit like the ST3000DM001 is a whole different thing entirely. :-) Yeah, there are times where the failure rate can rise so high it threatens the d
by brianwski 6y ago
> Guessing shit like the ST3000DM001 is a whole different thing entirely.
:-) Yeah, there are times where the failure rate can rise so high it threatens the data durability. The WORST is when failures are time correlated. Let's say the same one capacitor dies on a particular model of drive after precisely 6 months of being powered up. So everything is all calm and happy and smooth in operations, and then our world starts going sideways 1,200 drives at a time (one "vault" - our minimum unit of deployment).
Internally we've talked some about staggering drive models and drive ages to make these moments less impactful. But at any one moment one drive model usually stands out at a good price point, and buying in bulk we get a little discount, so this hasn't come to be.
- benlivengood 6y ago> Internally we've talked some about staggering drive models and drive ages to make these moments less impactful. But at any one moment one drive model usually stands out at a good price point, and buying in bulk we get a little discount, so this hasn't come to be. I don't know what your software architecture looks like right now (after reading the 2019 Vault post) but at some point it probably makes sense to move file shard location to a metadata layer to support more flexible layouts to work around failure domains (age, manufacturer, network switch, rack, power bus, physical location, etc.), reduce hotspot disks, and allow flexible hardware maintenance. Durability and reliability can be improved with two levels of RS codes as well; low level (M of N) codes for bit rot and failed drives and a higher level of (M2 of N2) codes across failure domains. It costs the same (N/M)*(N2/M2) storage as a larger (M*M2 of N*N2) code but you can use faster codes and larger N on the (N,M) layer (e.g. sse-accelerated RAID6) and slower, larger codes across transient failure domains under the assumption that you'll rarely need to reconstruct from the top-level parity, and any 2nd-level shards that do need to be reconstructed will be using data from a much larger number of drives than N2 to reduce hotspots. This also lets you rewrite lost shards immediately without physical drive replacement which reduces the number of parities required for a given durability level. This paper does something similar with product codes: http://pages.cs.wisc.edu/~msaxena/new/papers/hacfs-fast15.pdf http://pages.cs.wisc.edu/~msaxena/new/papers/hacfs-fast15.pd...