4 ms·
Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting
by jaketay 10y ago
Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing to take responsibility for this failure.
Whereas Google was recently down for less than 18 minutes. A VP at Google sent an email advising all affected customers, posted continuous updates to their status page, sent a further apology email at the conclusion, posted a service credit exceeding the SLA to all customers in the zone (without forcing customers to chase this themselves with billing) and lastly wrote one of the most well written post mortems I've ever seen. AWS has much to learn from Google about how to handle outages properly.
- Fenntrek 10y agoTo be fair, that was a total outage of every region on the planet in google cloud platform's 18 minute outage case. Compared to the sole region ap-southeast-2 (for you) in AWS's case. (though it was a longer outage). That being said, I 1000% agree that Google cloud platform's response to fix issues, postmortems and actions to make things right are top notch.
- ben_jones 10y agoI wonder if due to the scale of AWS and certain AWS customers, if AWS signs a post mortem with -Fred would a very large AWS customer have the pull to say "Amazon, fire Fred"? Just curious if anonymous postmortems are company policy at certain places and why that might be.
- deleted 10y ago[deleted]
- beachstartup 10y agopeople vote with their dollars and their votes are overwhelmingly telling amazon they're doing a great job.
- brazzledazzle 10y agoFor the record I agree with you but existing customers who are heavily invested in AWS would find it difficult to vote with their dollars.
- atonse 10y agoAgreed. I have a couple of clients that have a pretty substantial AWS spend, but the cost of switching to Azure is too high compared to the difference in offerings. You don't want to spend tens of thousands of dollars of developer time and risk switching datacenters for a small improvement.
- beachstartup 10y agoyeah, it works out great for amazon. less work, more money, what's not to love.
- jaketay 10y agoIt's difficult to change when you have reserved instances. A business decision yesterday might not be the best decision today. That being said, I think overall AWS is very good but Google is definitely starting to create some real competition which is positive for everyone.
- ejdyksen 10y ago> We used multiple availability zones and it didn't prevent downtime at all. Can you explain this a little more? Amazon says this only affected one AZ, and they specifically note: For this event, customers that were running their applications across multiple Availability Zones in the Region were able to maintain availability throughout the event.
- bigiain 10y ago+1. Apart from one internal project which mistakenly had all it's app server instances in -2b (ooops!) - all my production mobile app backends are spread across the 3 Sydney AZs. That's a few dozen EC2 app servers across about 15 projects. My monitoring reported a worst case of 57 seconds of degraded connectivity - which was an instance in -2b going offline and the ELB not taking it out of the rotation very quickly, the app running on that had interruption, but only while waiting for the timeouts. Crashlytics and GA crash reporting didn't bat an eyelid... I had under 70 users active at the time, 1/3rd of them may have seen a minute or less of loading spinner if they'd fired of a UI blocking api call during those 57 seconds. I'm not looking _super_ closely, but nothing I'm monitoring apart from EC2 - like RDS, S3, ELB, SNS - showed _any_ glitches (I'd _probably_ have caught even single digit second problems for _some_ of that...) I'm actually quite happy with how everything went - we don't go to any particular heroic lengths to ensure HA or uptime, we just follow recommended best practice, and at least in this outage, that worked out fine for us (except for that project where all the app servers were in -2b, and I'm happy to wear that as our fuckup)
- deleted 10y ago[deleted]
- jread 10y agoI independently monitor availability of 150 public cloud services and only observed 1.73 hours downtime for this event. This is the first EC2 outage I've observed in any region for over 6 months. According to my stats, in 2015 EC2 was highly available with 6 of 9 regions (including every US and EU region) having no outages, and total average service availability of 99.998% (78% of downtime in sa-east-1). I haven't observed a single outage in us-west-2 or eu-west-1 in nearly 3 years. This compared to 16-33 minutes of downtime in every region for GCE in 2015, and total average availability of 99.995%. Additionally, since 2013 I've never observed a global EC2 outage like the 4/11/2016 GCE event. https://cloudharmony.com/status-for-aws https://cloudharmony.com/status-for-aws
- brazzledazzle 10y agoMost of the criticism seemed to be centered around communication/PR so statistics don't really address that. Definitely important when considering a provider though. While the numbers are nice, if some outages only impact a subset of customers and your monitoring accounts aren't one of them it's hard to determine how good your monitoring data really is. If he was impacted by a 12 hour outage and you only show ~2 hours that's a really significant difference. I guess it really depends on how you monitor and how comprehensive it is. Do you monitor from multiple ISPs on different network paths in multiple regions/countries? Do you monitor each of the services under different load conditions and monitor multiple accounts? Sometimes "up" only tells part of the story.
- jread 10y ago2 hours aligns with the postmortem, 12 hours does not. The monitoring is based on a sampling of instances running in each service and region. Outages are verified from multiple network paths. While not comprehensive, over the years this has been generally accurate because most outage events have been network or power related, thus impacting most or all instances in the affected data center. There have been some isolated hardware failures, but they are rare.
- brazzledazzle 10y ago
- plandis 10y agoDo you have proof? Your claims directly contradict what was written in their post mortem both in terms of time and scope. AWS might be using nice wording but I can't see them lying that much about the scope of the incident.