4 ms·
Disclaimer: I work at Backblaze. > I wonder why Western Digital is almost absent, does anyone know why? Most of the time the answer comes down to price/GByte.
by brianwski 6y ago
Disclaimer: I work at Backblaze.
> I wonder why Western Digital is almost absent, does anyone know why?
Most of the time the answer comes down to price/GByte. But it isn't QUITE as simple as that.
Backblaze tries to optimize for total cost most of the time. That isn't just the cost of the drive, a drive that is twice as large in storage still takes the same identical amount of rack space and often the same electricity as the drive that is half the storage. This means that we have a spreadsheet and calculate what the total cost over a 5 year expected lifespan will turn out to be. So for example, even if the drive that is twice as large costs MORE than twice as much it can still make sense to purchase it.
As to failure rates, Backblaze essentially doesn't care what the failure rate of a drive is, other than to factor that into the spreadsheet. If we think one particular drive fails 2% more of the time, we still buy it if it is 2% cheaper, make sense?
So that's the answer most of the time, although Backblaze is always making sure we have alternatives, so we're willing to purchase a small number of pretty much anybody's drives of pretty much any size in order to "qualify" them. It means we run one pod of 60 of them for a month or two, then we run a full vault of 1,200 of that drive type for a month or two, just in case a good deal floats by where we can buy a few thousand of that type of drive. We have some confidence they will work.
- Hamuko 6y ago>As to failure rates, Backblaze essentially doesn't care what the failure rate of a drive is, other than to factor that into the spreadsheet. Guessing shit like the ST3000DM001 is a whole different thing entirely.
- brianwski 6y ago> Guessing shit like the ST3000DM001 is a whole different thing entirely. :-) Yeah, there are times where the failure rate can rise so high it threatens the data durability. The WORST is when failures are time correlated. Let's say the same one capacitor dies on a particular model of drive after precisely 6 months of being powered up. So everything is all calm and happy and smooth in operations, and then our world starts going sideways 1,200 drives at a time (one "vault" - our minimum unit of deployment). Internally we've talked some about staggering drive models and drive ages to make these moments less impactful. But at any one moment one drive model usually stands out at a good price point, and buying in bulk we get a little discount, so this hasn't come to be.
- benlivengood 6y ago> Internally we've talked some about staggering drive models and drive ages to make these moments less impactful. But at any one moment one drive model usually stands out at a good price point, and buying in bulk we get a little discount, so this hasn't come to be. I don't know what your software architecture looks like right now (after reading the 2019 Vault post) but at some point it probably makes sense to move file shard location to a metadata layer to support more flexible layouts to work around failure domains (age, manufacturer, network switch, rack, power bus, physical location, etc.), reduce hotspot disks, and allow flexible hardware maintenance. Durability and reliability can be improved with two levels of RS codes as well; low level (M of N) codes for bit rot and failed drives and a higher level of (M2 of N2) codes across failure domains. It costs the same (N/M)*(N2/M2) storage as a larger (M*M2 of N*N2) code but you can use faster codes and larger N on the (N,M) layer (e.g. sse-accelerated RAID6) and slower, larger codes across transient failure domains under the assumption that you'll rarely need to reconstruct from the top-level parity, and any 2nd-level shards that do need to be reconstructed will be using data from a much larger number of drives than N2 to reduce hotspots. This also lets you rewrite lost shards immediately without physical drive replacement which reduces the number of parities required for a given durability level. This paper does something similar with product codes: http://pages.cs.wisc.edu/~msaxena/new/papers/hacfs-fast15.pdf http://pages.cs.wisc.edu/~msaxena/new/papers/hacfs-fast15.pd...
- louwrentius 6y agoThank you for this very elaborate and detailed answer! Disclaimer: (I’m a customer)
- hinkley 6y agoIs it safe to say that Backblaze essentially has an O(logn) algorithm for labor due to drive installation and maintenance, so up front costs and opportunity costs due to capacity weigh heavier in the equation? The rest of us don’t have that, so a single disk loss can ruin a whole Saturday. Which is why we appreciate that you guys post the numbers as a public service/goodwill generator.
- brianwski 6y ago> algorithm for labor due to drive installation and maintenance ... the rest of us don't have that so a single disk loss can ruin Saturday TOTALLY true. We staff our datacenters with our own datacenter technicians (Backblaze employees) 7 days a week. When they arrive in the morning the first thing they do is replace any drives that failed during the night. The last thing they do before going home is replacing the drives that failed during the day so the fleet is "whole". Backblaze currently runs at 17 + 3. 17 data drives with 3 calculated parity drives, so we can lose ANY THREE drives out of a "tome" of 20 drives. Each of the 20 drives in one tome is in a different rack in the datacenter. You can read a little more about that in this blog post: https://www.backblaze.com/blog/vault-cloud-storage-architecture/ https://www.backblaze.com/blog/vault-cloud-storage-architect... So if 1 drive fails at night in one 20 drive tome we don't wake anybody up, and it's business as usual. That's totally normal, and the drive is replaced at around 8am. However, if 2 drives fail in one tome pagers start going off and employees wake up and start driving towards the datacenter to replace the drives. With 2 drives down we ALSO automatically stop writing new data to that particular tome (but customers can still read files from that tome), because we have notice less drive activity can lighten failure rates. In the VERY unusual situation that 3 drives are down in one tome every single tech ops and datacenter tech and engineer at Backblaze is awake and working on THAT problem until the tome comes back from the brink. We do NOT like being in that position. In that situation we turn off all "cleanup jobs" on that vault to lighten load. The cleanup jobs are the things that are running around deleting files that customers no longer need, like if they age out due to lifecycle rules, etc. The only exceptions to our datacenters having dedicated staff working 7 days a week are if a particular datacenter is small or just coming online. In that case we lean on "remote hands" to replace drives on weekends. That's more expensive per drive, but it isn't worth employing datacenter technicians that are just hanging out all day Saturday and Sunday bored out of their minds - instead we just pay the bill for remote hands.