7 ms·
> I haven't heard of people having problems [with S3's Durability] but equally: I've never seen these claims tested. I am at least a bit curious about these cla
by breckognize 3y ago
> I haven't heard of people having problems [with S3's Durability] but equally: I've never seen these claims tested. I am at least a bit curious about these claims.
Believe the hype. S3's durability is industry leading and traditional file systems don't compare. It's not just the software - it's the physical infrastructure and safety culture.
AWS' availability zone isolation is better than the other cloud providers. When I worked at S3, customers would beat us up over pricing compared to GCP blob storage, but the comparison was unfair because Google would store your data in the same building (or maybe different rooms of the same building) - not with the separation AWS did.
The entire organization was unbelievably paranoid about data integrity (checksum all the things) and bigger events like natural disasters. S3 even operates at a scale where we could detect "bitrot" - random bit flips caused by gamma rays hitting a hard drive platter (roughly one per second across trillions of objects iirc). We even measured failure rates by hard drive vendor/vintage to minimize the chance of data loss if a batch of disks went bad.
I wouldn't store critical data anywhere else.
Source: I wrote the S3 placement system.
- tracerbulletx 3y agoMy first job was at a startup in 2012 where I was expected to build things at a scale way over what I really had the experience to do. Anyways the best choice I ever made was using RDS and S3 (and django).
- notso411 3y ago[dead]
- supriyo-biswas 3y agoChecksumming the data is not based out of paranoia but simply as a result of having to detect which blocks are unusable in order to run the Reed-Solomon algorithm. I'd also assume that a sufficient number of these corruption events are used as a signal to "heal" the system by migrating the individual data blocks onto different machines. Overall, I'd say the things that you mentioned are pretty typical of a storage system, and are not at all specific to S3 :)
- catlifeonmars 3y agoThe S3 checksum feature applies to the objects, so that’s entirely orthogonal to erasure codes. Unless you know something I don’t and SHA256 has commutative properties. You’d still need to compute the object hash independent of any blocks. Source: https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/checki...
- benlivengood 3y agoIt's not entirely orthogonal; RAID5 plus stripe-level CRC (or better) can reliably correct bitrot at any single position in a stripe whereas RAID5 alone can only report an error. My guess is that S3 and other large object stores have the equivalent of stripe-level checksums for this purpose.
- catlifeonmars 3y agoI’m positive something like this is the case. Yet it’s entirely orthogonal to the object hash in the user facing feature, which would need to be computed separately.
- benlivengood 3y agoFor append-only or write-once objects or for BLAKE-3 and other fully parallelizable hashes it's possible to store the intermediate hash function state with each chunk or stripe so that the final bytes of the data, once the hash is finished, yield the user-facing checksum as well.
- staunch 3y ago> Believe the hype. I'd rather believe the test results. Is there a neutral third-party that has validated S3's durability/integrity/consistency? Something as rigorous as Jepsen? It'd be really neat if someone compared all the S3 compatible cloud storage systems in a really rigorous way. I'm sure we'd discover that there are huge scary problems. Or maybe someone already has?
- rsync 3y ago"AWS' availability zone isolation is better than the other cloud providers." Not better than all of them. A geo-redundant rsync.net account exists in two different states (or countries) - for instance, primary in Fremont[1] and secondary in Denver. "S3 even operates at a scale where we could detect "bitrot"" That is not a function of scale. My personal server running ZFS detects bitrot just fine - and the scale involved is tiny. [1] he.net headquarters
- Helmut10001 3y agoAgree. > S3 even operates at a scale where we could detect "bitrot" - random bit flips caused by gamma rays hitting a hard drive platter (roughly one per second across trillions of objects iirc). I would expect any cloud provider to be able to detect bitrot these days.
- senderista 3y agoI think the point the OP was trying to make is that they regularly detected bitrot due to their scale, not that they were merely capable of doing so.
- Helmut10001 3y agoAh, thank you. This makes more sense. And I think I remember reading about it once. Apologies for the misinterpretation!
- pclmulqdq 3y agoEveryone with significant scale and decent software regularly detects bitrot.
- breckognize 3y agoBacking up across two different regions is possible for any provider with two "regions" but requires either doubling your storage footprint or accepting a latency hit because you have to make a roundtrip from Fremont to Denver. The neat thing about AWS' AZ architecture is that it's a sweet spot in the middle. They're far enough apart for good isolation, which provides durability and availability, but close enough that the network round trip time is negligible compared to the disk seek. Re: bit rot, I mean the frequency of events. If you've got a few disks, you may see one flip every couple years. They happen frequently enough in S3 that you can have expectations about the arrival rate and alarm when that deviates from expectations.
- deleted 3y ago[deleted]
- medler 3y ago> customers would beat us up over pricing compared to GCP blob storage, but the comparison was unfair because Google would store your data in the same building I don’t think this is true. Per the Google Cloud Storage docs, data is replicated across multiple zones, and each zone maps to a different cluster. https://cloud.google.com/compute/docs/regions-zones/zone-virtualization https://cloud.google.com/compute/docs/regions-zones/zone-vir...
- singron 3y agoGoogle puts multiple clusters in a single building.
- yencabulator 3y agoZones are about correlated power and networking failures. Regions are about disasters. If you want multiple regions, Google can of course do that too: https://cloud.google.com/storage/docs/locations#considerations https://cloud.google.com/storage/docs/locations#consideratio...
- treflop 3y agoWhat’s your experience like at other storage outfits? I only ask because your post is a bit like singing praises for Cinnabon that they make their own dough. The things that you mentioned are standard storage company activities. Checksum-all-the-things is a basic feature of a lot of file systems. If you can already set up your home computer to detect bitrot and alert you, you can bet big storage vendors do it. Keeping track of hard drive failure rates by vendor is normal. Storage companies publicly publish their own reports. The tiny 6-person IT operation I was in had a spreadsheet. Hell, I toured a friend’s friend’s major data center last year and he managed to find time to talk hard drive vendors. Now you. I get it — y’all make spreadsheets. There are a lot of smart people working on storage outside AWS and long before AWS existed.
- pclmulqdq 3y agoWhen I worked at Google in storage, we had our own figures of merit that showed that we were the best and Amazon's durability was trash in comparison to us. As far as I can tell, every cloud provider's object store is too durable to actually measure ("14 9's"), and it's not a problem.
- breckognize 3y ago9's are overblown. When cloud providers report that, they're really saying "Assuming random hard drive failure at the rates we've historically measured and how we quickly we detect and fix those failures, what's the mean time to data loss". But that's burying the lede. By far the greatest risks to a file's durability are: 1. Bugs (which aren't captured by a durability model). This is mitigated by deploying slowly and having good isolation between regions. 2. An act of God that wipes out a facility. The point of my comment was that it's not just about checksums. That's table stakes. The main driver of data loss for storage organizations with competent software is safety culture and physical infrastructure. My experience was that S3's safety culture is outstanding. In terms of physical separation and how "solid" the AZs are, AWS is overbuilt compared to the other players.
- pclmulqdq 3y ago
- loeg 3y agoNot a public cloud, but storage at Facebook is similar in terms of physical infrastructure, safety culture, and scale.
- spintin 3y agoCorrect me if I'm wrong but bitrot only affects spinning rust since NAND uses ECC? If you see this I wonder if S3 is planning on adding hardlinks?
- Veserv 3y agoBut they asked if the claims were audited by a unbiased third party. Are there such audits? Alternatively, AWS does publicly provide legally binding availability guarantees, but I have never seen any prominently displayed legally binding durability guarantees. Are these published somewhere less prominently?
- cyberax 3y ago> Alternatively, AWS does publicly provide legally binding availability guarantees, but I have never seen any prominently displayed legally binding durability guarantees. Are these published somewhere less prominently? It's listed prominently in the public docs: https://aws.amazon.com/s3/storage-classes/ https://aws.amazon.com/s3/storage-classes/
- Veserv 3y agoI read that page and it does not provide any contractual durability guarantees as far as I can see. It provides "designed for availability" and then contractual availability SLA guarantees. It provides "designed for durability", but presents no contractual durability guarantee as far as I can see. Given that their lawyers clearly indicate that "designed for availability" is not what they are contractually obligated to provide, only the letter of the SLA does that; "designed for durability" is similarly a marketing statement that does not incur any contractual obligations. Is there some specific statement in that document that I am missing which indicates that data durability is not fully at their convenience?
- AdamN 3y agoSLAs are more of a financial construct than anything else. Once the payback cost of missing an SLA is built into the contract then it just becomes a conversation about money. I've been at plenty of shops that obviously tried to hit the SLA but if it was missed it just became a financial issue which helped smooth over what otherwise might have been a trust buster. I would never ever think of an SLA as anything more than a financial commitment - if you think more of it you'll eventually be in a world of hurt.
- duckman1 3y ago[dead]
- notso411 3y ago[dead]
- chupasaurus 3y ago> and bigger events like natural disasters Outdated anecdata: I've worked for a company that lost some parts of buckets after the lightning strike incident in 2011, which bumped the paranoia quite a bit. AFAIK same thing couldn't happen for more than a decade.
- simonebrunozzi 3y agoI also worked at AWS, but not in the S3 team. However, I was Tech Evangelist and met with literally thousands of customers over my 6 years tenure. S3 was one of the hottest topics, but I got a sense of how good and robust it was directly from these customers. What you say resonates really well with me, and what I've heard during these years.
- zooq_ai 3y agoGoogle discovered random bit flips caused by gamma rays.