5 ms·
Heroku Postmortem of June 29th Incident
- aaronsw 14y agoThe "What we are doing" section seems pretty weak. The only substantive thing they say is "we have produced new tools which enable us to more expediently relocate database services from a failed availability zone." How exactly are they planning to deal with the larger Cedar difficulties? Are they going to eliminate their dependence on ELBs? Go multi-region? Developers need to know this to decide whether to continue with Heroku or build their own platform.
- rdl 14y agoCompare/contrast to https://status.heroku.com/incidents/151 https://status.heroku.com/incidents/151 where they did talk about moving to multiregion, etc., but then never executed on it, which is probably worse. Maybe they will underpromise and overdeliver. Personally, I'd shitcan EBS to the extent possible, RDS I never would have used, and it looks like ELB is non-viable as well. If you're a large site (particularly a PaaS) on AWS and care about availability, you need to have spare capacity in your region (using RIs, like Netflix does) to cover when a single AZ disappears, and your own external to AWS load balancing (not dns based), with your own per-AZ subsidiary load balancers (nginx or whatever) running within EC2. You need a robust database layer, ideally multi-region or AWS+nonAWS, but that's more site specific. Going multiregion is the next step, and the above is an essential part of getting to that point.
- jscottmiller 14y ago> Approximately 30% of our EC2 instances, which were responsible for running applications, databases and supporting infrastructure (including some components specific to the Bamboo stack), went offline Combined with the incident report from amazon, does this mean that 30% of Heroku instances were in a single availability zone? That would be troubling.
- dangrossman 14y agoWhy is that troubling? There are currently 5 AZs in the region (there were fewer when they designed their service), so the best they could possibly do is 20% of their instances in each. Losing 20% of their instances instead of 30% isn't much of an improvement. What they need is failover between regions/clouds and external monitoring, not a slightly better allocation of instances between AWS AZs in one region.
- jscottmiller 14y agoI don't use Heroku, but I would have assumed that they exist in more than one AWS region.
- dangrossman 14y agoThey don't.
- jbaudanza 14y agoAccording to this tweet, they are in 3 AZs. https://twitter.com/heroku/status/219937314749677568 https://twitter.com/heroku/status/219937314749677568
- dangrossman 14y ago3 AZs all in the US-East region. That's consistent with what we're talking about.
- goronbjorn 14y agoMy questions are: - Why aren't they committing to using geographically dispersed AWS instances? - Why aren't they leveraging Salesforce's infrastructure at all?
- rdl 14y agoYou may not have noticed the 2x multi-hour outages Salesforce has had in the past 2 weeks, at least one due to "power problems", the more recent of which took down their fucking status site too.
- goronbjorn 14y agoMy point wasn't that Salesforce's infrastructure is reliable (although they arguably have more pressure than most companies to be so), just that there is a separate resource they could be utilizing as a backup/in addition to AWS that they currently are not using.
- antoko 14y agoPresumably if they were having to deal with a power outage you could reliably assume that that fucking status was set to "not fucking"
- dangrossman 14y agoCommitting to geographically dispersed anything is a hard problem. How will apps perform when their front-end is in Virginia and their back-end is in California? How will fail-over work when it involves moving a 100GB database from Dallas to Washington, instead of between racks in the same building over a private gigabit network?
- goronbjorn 14y agoI agree that it's a difficult problem, but it's Heroku's obligation to solve those problems. You can't position yourself like this: > Heroku takes full responsibility for your app's health, keeping it up and running through thick and thin and not commit to solving hard problems like geographically dispersed EC2 instances. Resource constraints, to be fair, aren't an acceptable excuse at this point. They aren't a floundering startup. They're a part of a multibillion dollar company.
- latch 14y agoRelying on AWS' API to mitigate an AWS failure seems dangerous. A fundamental catch-22 with on-demand provisioning/configuration.
- paulsutter 14y agoOne subtle but important reason to use cross-region failover is that the network latency between the regions can prevent many casual or accidental dependencies between instances (if you configure instances in two regions to use the same database server, latency can cause the distant region to perform poorly). This is why it's really hard to get cross region failover to work. Because you really need to make them independent.
- rhizome 14y agoYou're describing a premature optimization. Performing poorly is better than not performing at all.
- adrianpike 14y agoOne of their suggestions is to have a follower of your DB and fall back to it. When they put the API in read-only mode, would I have been able to promote any followers?