4 ms·
I work at Backblaze. > how many physical locations to they run? Two separate datacenters in the Sacramento (California) region, and one in Phoenix (Arizona).
by brianwski 8y ago
I work at Backblaze.
> how many physical locations to they run?
Two separate datacenters in the Sacramento (California) region, and one in Phoenix (Arizona). We are trying to open a European (Netherlands) datacenter this month or next month.
However, unless you take explicit action to copy your data to two datacenters, any one file (or piece of file) is in exactly one datacenter. We believe the data to be extremely DURABLE (survive), but if your strategy is to "host content" in a highly available fashion where people will die if your content is offline for an hour, we recommend you use two different providers with some sort of fail over. Another alternative is to use a CDN (Content Delivery Network) in conjunction with Backblaze. You can find out more info here: https://www.backblaze.com/b2/solutions/content-delivery.html https://www.backblaze.com/b2/solutions/content-delivery.html
For backups, Backblaze advocates for a 3-2-1 backup strategy. https://www.backblaze.com/blog/the-3-2-1-backup-strategy/ https://www.backblaze.com/blog/the-3-2-1-backup-strategy/ This is where you keep 3 copies of your data, 2 on site, and 1 in the cloud.
The exact system for how we achieve high durability is described in this blog past: https://www.backblaze.com/blog/vault-cloud-storage-architecture/ https://www.backblaze.com/blog/vault-cloud-storage-architect... where any one file is striped across 20 separate computers in 20 separate locations in the one datacenter, where we can entirely lose any 3 computers and the data is completely fine and available.
We are COMPLETELY transparent on how we calculate the durability, we do the math (including the assumptions) in this blog post: https://www.backblaze.com/blog/cloud-storage-durability/ https://www.backblaze.com/blog/cloud-storage-durability/
- sp332 8y agoTo be a little more specific, the "2" in 3-2-1 is for two different media types. Hard drive + tape, for example. [Edit: ok it doesn't have to be "types". But two different media - don't put all your backups on one disk!]
- atYevP 8y agoYev from Backblaze here -> Or Hard Drive (Internal) + Hard Drive (External) - that's what we typically see!
- chillaxtian 8y agoNo disaster recovery? :/
- toomuchtodo 8y agoSame as every other storage provider's default/basic storage offering. If you want georedundancy, you will need to build it. EDIT: Apparently GCS has this feature built in. Did not know, very cool!
- chrisseaton 8y agoI thought basic products like S3 provided cross-region replication, which gives georedundancy? But anyway why should I have to build it on top of the provider's offering - why wouldn't they provide georedundancy for me? It seems like a truly basic thing to expect for a backup solution? But I'm not an expert in this area.
- deleted 8y ago[deleted]
- siculars 8y agoGood Cloud Storage (GCS) has this functionality out of the box. https://cloud.google.com/storage/docs/locations https://cloud.google.com/storage/docs/locations /I work for Google/
- luhn 8y agoS3 replicates across available zones, meaning copies in multiple DCs but in the same general area. You can setup a bucket to replicate a bucket in another region, at double the storage costs plus bandwidth charges.
- manigandham 8y agoAll clouds have options to do multi-regional storage. GCS has multiregional class. Azure has GRS class. AWS has cross-region replication that can be added to a bucket.
- luhn 8y agoLast I checked, the backup service exclusively used the Sacramento DC. Has this changed? Being in the Sacramento area myself, I'd be much more comfortable if my offsite backup was more than a couple miles away.
- Johnny555 8y agoWe believe the data to be extremely DURABLE (survive), but if your strategy is to "host content" in a highly available fashion where people will die if your content is offline for an hour, we recommend you use two different providers with some sort of fail over Do your durability metrics take datacenter failure or human error into account? Datacenter failures are rare, but they do happen and can cause data loss if all of your data is in that data center. Likewise, human error can cause cascading failures across a datacenter (or beyond) if there are no firewalls between zones that prevent a single person/command/software update from affecting all copies of the data.
- manigandham 8y agoThat depends on what you mean by failure. Are you talking about a data center failing because every single machine inside blew up? Otherwise the common failures like DC power outage or network drops are about availability rather than durability. The data is stills safe on multiple drives.
- Johnny555 8y agoI'm talking about the kind of failure that hit a Microsoft Azure data center: https://www.datacenterknowledge.com/microsoft/azure-outage-proves-hard-way-availability-zones-are-good-idea https://www.datacenterknowledge.com/microsoft/azure-outage-p... “ but in this instance, temperatures increased so quickly in parts of the data center that some hardware was damaged before it could shut down" ... "A significant number of storage servers were damaged, as well as a small number of network devices and power units.”
- manigandham 8y agoYea if that damaged all the storage servers containing your data then there would be data loss.
- Johnny555 8y agoThat's why I asked if that 99.999999999% durability number includes datacenter loss. It's an unlikely failure mode but is it .00000000001% unlikely? I don't know. Given the fact that Azure lost a datacenter with this failure mode, I don't think it's in the "likelihood of an asteroid destroying Earth within a million years" ballpark. Their durability page doesn't really clear it up, they say "Because at these probability levels, it’s far more likely that ... Earthquakes / floods / pests / or other events" known as “Acts of God” destroy multiple data centers". But from the post above: "any one file (or piece of file) is in exactly one datacenter." So it doesn't take multiple datacenter failures to lose data, just one unless you explicitly copy your data to multiple datacenters.