4 ms·
Would be interesting to also have metrics on failure per TB storage.
by alinde 8y ago
Would be interesting to also have metrics on failure per TB storage.
- theandrewbailey 8y agoI'm not sure how that would be useful, since terabytes don't fail. When a drive fails, it's effectively a brick with no terabytes.
- alinde 8y agoI was thinking as one failure of a 100TB disk has a very different impact of 10 failures of 1TB disks. It'd give some idea on how much data is lost due failures, no?
- atYevP 8y agoYev from Backblaze here -> Not sure if you'd get that metric from that data. We use Reed-Solomon erasure coding (https://www.backblaze.com/blog/reed-solomon/ https://www.backblaze.com/blog/reed-solomon/) to make sure that data is "rebuilt" should we lose drives (which happens all the time).
- deleted 8y ago[deleted]
- oliveshell 8y agoI suppose, but there’s no such thing as a single HDD that stores 100TB. The biggest you can get currently are (I believe) 14TB helium-filled drives.
- jsgo 8y agoMy guess is they meant 10TB as it would be a more "equal" comparison: 1 10TB drive 10 1TB drives
- oliveshell 8y agoThat makes way more sense. Didn’t think before posting!
- brianwski 8y agoDisclaimer: I work at Backblaze. > When a drive fails, it's effectively a brick with no terabytes. Interesting factoid: that isn't always true. What you describe is actually the CLEANEST type of failure, the drive suddenly becomes a brick. We replace the drive and rebuild it from parity. A way more interesting failure is when disk blocks start going bad at an unacceptable rate. Backblaze splits your data across 20 different hard drives in 20 different machines in our datacenter. The sub-parts we call "shards", a shard sits on one disk. Each shard has a SHA-1 checksum, so we know if each shard has been corrupted. If an individual shard is missing or corrupted, we know it needs to be rebuilt from parity. So when a drive is HALF-FAILED, we even have a procedure to pull the drive out, and then opportunistically copy whatever files we can recover onto a new drive, then put the new drive back into production. Any files we recover where they are in the correct filesystem location and their SHA1 says they have not been corrupted speeds up the rebuild. The reason the speed of rebuild is important is the whole concept of 11 or 12 "nines" of durability. We can't have more than 3 drives fail in any one group of 20 drives, and the faster the rebuild time, the less likely for 4 simultaneous failures. It plugs into the formulas in this blog post we did about durability: https://www.backblaze.com/blog/cloud-storage-durability/ https://www.backblaze.com/blog/cloud-storage-durability/
- zepearl 8y ago(thanks a lot - all extremely interesting) >> We can't have more than 3 drives fail in any one group of 20 drives... Wow, for me, subjectively, an low threshold - and I underderstand that each drive being hosted on a different machine protects you as well from a machine/controller failure (happened to me twice with the controller - both times it was very hard to diagnose and the experience in general has been terrible). Do you have as well "backups"? Or is that in the hands of the customers/users?
- brianwski 8y ago> Do you have as well "backups"? Or is that in the hands of the customers/users? If you store data in Backblaze, there is no "backup" of that data. If Backblaze ever lost 4 drives simultaneously and could not recover the data, the customer would lose data. This is much like Amazon S3. In general, we recommend a 3-2-1 backup strategy where there are 3 copies of the data, at least 2 copies on your site, and 1 copy in the cloud. You can read about that philosophy in our blog post here: https://www.backblaze.com/blog/the-3-2-1-backup-strategy/ https://www.backblaze.com/blog/the-3-2-1-backup-strategy/
- klodolph 8y agoI’m not sure that it’s especially useful to measure that way, which is why they wouldn’t report it. The chance that a given GB of data is on a failed disk is equal to the disk failure rate, regardless of disk size (>1GB). For large deployments, the concern is between failure rates and the amount of time it takes to rebuild data from a failed disk. For small deployments, my main concerns are whether disk failure takes a machine or volume out, causing availability problems. I’m trying to figure where failures per GB would be how you would choose, what scenario we’re you thinking of?