7 ms·
9's are overblown. When cloud providers report that, they're really saying "Assuming random hard drive failure at the rates we've historically measured and how
by breckognize 3y ago
9's are overblown. When cloud providers report that, they're really saying "Assuming random hard drive failure at the rates we've historically measured and how we quickly we detect and fix those failures, what's the mean time to data loss".
But that's burying the lede. By far the greatest risks to a file's durability are:
1. Bugs (which aren't captured by a durability model). This is mitigated by deploying slowly and having good isolation between regions.
2. An act of God that wipes out a facility.
The point of my comment was that it's not just about checksums. That's table stakes. The main driver of data loss for storage organizations with competent software is safety culture and physical infrastructure.
My experience was that S3's safety culture is outstanding. In terms of physical separation and how "solid" the AZs are, AWS is overbuilt compared to the other players.
- pclmulqdq 3y agoThat was not how we treated the 9's at Google. Those had been tested through natural experiments (disasters). I was not at Google for the Clichy fire, but it wasn't the first datacenter fire Google experienced. I think your information about Google's data placement may be incorrect, or you may be mapping AWS concepts onto Google internal infrastructure in the wrong way.
- fsociety 3y agoI would not lose sleep over storing data on GCS, but have heard from several Google Cloud folks that their concept of zones is a mirage at best.
- pclmulqdq 3y agoYeah, that's definitely true. Google sort of mapped an AWS concept onto its own cluster splits. However, there are enough regional-scale outages at all the major clouds that I don't personally place much stock in the idea of zones to begin with. The only way to get close to true 24/7 five-9's uptime with clouds is to be multi-region (and preferably multi-cloud).
- callalex 3y agoI have experienced many outages that were contained to a specific availability zone in AWS, from power failures to flooding to cable cuts. You are correct that 5 9’s still requires multi-region though.
- kevincox 3y agoI think also Google as a whole has pretty good diversity. But Cloud customers demanded regions in big population centers and smaller countries where Google traditionally avoided due to cost reasons. This lead to less redundant sites that were often owned and/or operated by third parties. So in the US and Europe you can probably trust GCP zones quite literally. But other regions (I have heard lots of rumours in the APAC) they may not be quite as diverse as they appear.
- throwaway2037 3y ago> big population centers and smaller countries Can we stop this dance on HN? Can you just name them, please?
- pclmulqdq 3y agoI think most Googlers actually don't know the specifics (I certainly don't know), and if they could, they probably couldn't tell you. It's sort of common knowledge that some of them are like this, but not exactly which ones.
- flaminHotSpeedo 3y agoSee: that fire in France that took down a whole region But on the other hand, GCP supports multi-region so that's not nearly as big of a deal as it would be if AWS zones were not sufficiently isolated
- breckognize 3y agoDo you mean Google included "acts of God" when computing 9's? That's definitely not right. 11 9's of durability means mean time to data loss of 100 billion years. Nothing on earth is 11 9's durable in the face of natural (or man-made) disasters. The earth is only 4.5 billion years old.
- pclmulqdq 3y agoNormally, companies store more than 1 byte of data, and the 9's (not just for data loss, for everything) are ensemble averages. By the way, I don't doubt that AWS has plenty of 9's by that metric - perhaps more than GCP.
- jftuga 3y agoIf I were to upload a 50kb object to S3 (standard tier), about how many unique physical copies would exist?
- cyberax 3y agoAt least 3.
- bigiain 3y agoAt least 3, in at least 3 seperate datacenters. According to https://nuclearsecrecy.com/nukemap/ https://nuclearsecrecy.com/nukemap/ - it'd take at least a 1 megaton warhead to take out two of the ap-southeast-2 datacenters, and over 10MT to take out 3. I suspect you'd need a lot less than that though, the 1MT warhead would probably take out enough outside-the-datacenter infrastructure to take the entire AZ offline. I don't care too much though, if someone's dropped a warhead that close to home I have other things to worry about than whether all the cat pictures and audit logs survive.
- tomcam 3y ago> if someone's dropped a warhead that close to home I have other things to worry about than whether all the cat pictures and audit logs survive. Speak for yourself. Many of us love our audit logs and show them to strangers whatever we can.
- kragen 3y ago0. user error (deletion or overwriting a file they regret later, possibly much later) -1. government, the historical cause of most data loss 1½. google canceling the product and deleting all the data, as they did with google+
- wubrr 3y ago9's are useful when they're backed by an actual SLA - like GCP Cloud Storage and AWS S3 availability SLAs. Neither one commits to any durability SLAs whatsoever so I wouldn't put any stock into the 'eleventy nine nines' durability claims.