4 ms·
If you're worried about a datacenter losing power, a multi-AZ strategy is fine. Hardware issues happen so often and with so little impact that we rarely hear ab
by mnutt 11y ago
If you're worried about a datacenter losing power, a multi-AZ strategy is fine. Hardware issues happen so often and with so little impact that we rarely hear about them. The issues that I'm primarily concerned about at this point are in software.
The S3 issue affected the entire region. A month back a BGP issue affected every one of our AZs in us-east-1. It wasn't even Amazon's fault, but a multi-AZ strategy would have done nothing for us.
It's kind of like Maslow's hierarchy of availability needs: at the bottom you deal with hardware failures, then networking failures, then regional configuration failures, and finally homogeneity issues where all of your machines fail at once because of a bug in hardware raid controllers or a leap second smear issue. It all depends on how much resilience you need and are willing to pay for.