3 ms·
I'm not sure how well this is going to work out in practice, but the concept of having multi-region failover could mean that much closer to almost entirely bull
by programminggeek 13y ago
I'm not sure how well this is going to work out in practice, but the concept of having multi-region failover could mean that much closer to almost entirely bulletproof infrastructure (if you are willing to spend the money).
- recuter 13y agoNot really. This is no different than running your own, say, Haproxy with a heartbeat between two boxes next to each other on the same rack or some such. That Amazon now with a bit of fiddling will fail over to another region if your ELB instance is down or because it detects errors is great - but hardly bullet proof. What if your ELB instance is ok but your app is returning garbage.. its still returning something, right? And ELB will happily forward that traffic without thinking anything is wrong.
- colmmacc 13y agoFull disclosure: Route 53 developer here. There is one interesting difference; when using DNS failover in combination with Latency Based Routing, Route 53 supports partition mode failures. For example; if a customer has endpoints available in both the AWS Sydney region and an AWS US region, and Australian international connectivity is impaired then users within Australia will still go to Sydney. At the same time, a user in New Zealand who would ordinarily go to Australia (as it's closer) may now find themselves served by US endpoints because reachability to Australia from New Zealand is impaired but reachability to the US is ok. It's a small part of the availability story, but one difference in how a DNS failover may handle an event. The "returning garbage" problem can be a hard one - both Route 53 and ELB can be configured to check a particular url for health-status, and it's important that that url's health be indicative of the overall stack's health, but it's definitely a challenge sometimes as an application operator to maintain a good "deep" check. For example, on my own personal Wordpress installation I check if the DB is reachable and answering before returning 200 from my status URL, but I've seen service owners do much much more comprehensive checks including inspecting some metrics, counting overall 500s, measuring response times, and so on.