3 ms·
They replicate the data by some amount within an availability zone using some form of forward error correction. They look at their machine/hard drive irrecovera
by posnet 6y ago
They replicate the data by some amount within an availability zone using some form of forward error correction. They look at their machine/hard drive irrecoverable failure rates within that facility and get 99.99% durability.
Then they replicate all of that again 3 times within a region across 3 different availability zones (again independent facilities). And they are then making the claim that they will never lose all 3 facilities irrecoverably at the same time. i.e (1 - 0.9999)^3 = 0.99999999999.
Simplified view, since they likely have to tweak the replication by object size since their claim is by object and not number of bytes. But, the answer is not magic, just redundancy.
And you can get very good at predicting failures when you have millions of hard drives.
It's part of the reason AWS never open a new region without at least 3 availability zones, otherwise they couldn't run s3 there.