7 ms·
EC2 or S3 showing red in any region literally requires personal approval of the CEO of AWS.
by ta20200710 5y ago
EC2 or S3 showing red in any region literally requires personal approval of the CEO of AWS.
- dia80 5y agoUnfortunately, errors don't require his approval...
- notreallyserio 5y agoIs this true or a joke? This sort of policy is how you destroy trust.
- bsedlm 5y agomaybe we gotta consider the publicly facing status pages as something other than a technical tool (e.g. marketing or PR or something like that, dunno)
- jedberg 5y agoFrom what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec. Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.
- notreallyserio 5y agoI wonder how well known this is. You'd think it would be hard to hire ethical engineers with such a scheme in place and yet they have tens of thousands.
- orangepurple 5y agoBeing dishonest about SLAs seems to bear zero cost in this case?
- jedberg 5y agoIt's not really dishonest though because there is nuance. Most everything in EC2 is still working it seems, just the console is down. So is it really down? It should probably be yellow but not red.
- dekhn 5y agoif you cannot access the control plane to create or destroy resources, it is down (partial availability). The jobs that are running are basically zombies.
- w0m 5y agoDepending the workload being run users may or may not notice. Should be Yellow at a minimum.
- jedberg 5y agoSeems like the API is still working and so is auto scaling. So they aren’t really zombies. Partial availability isn’t the same as no availability.
- electroly 5y agoThe API is NOT working -- it may not have been listed on the service health dashboard when you posted that, but it is now. We haven't been able to launch an instance at all, and we are continuously trying. We can't even start existing instances.
- dekhn 5y agoI'm right in the middle of an AWS-run training and we literally can't run the exercises because of this. let me repeat that: my AWS trainign that is run by AWS that I pay AWS for isn't working, because AWS is having control plane (or other) issues. This is several hours after the initial incident. We're doing training in us-west-2, but the identity service and other components run in us-east-1.
- dekhn 5y agoSure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean that large corps with influence should be talking to AWS to ensure that status information corresponds with actual service outages.
- emodendroket 5y agoI have no inside knowledge or anything but it seems like there are a lot of scenarios with degraded performance where people could argue about whether it really constitutes an outage.
- dekhn 5y agoYep. I was an SRE who worked at Google and also launched a product on Google Cloud. We had these arguments all the time, and the contract language often provides a way for the provider to weasel out.
- dilyevsky 5y agoOne time gcp argued that since they did return 404s on gcs for a few hours that wasn’t an uptime/latency sla violation so we were not entitled to refund (tho they refunded us anyway)
- Enginerrrd 5y agoMan, between costs and shenanigans like this, why don't more companies self-host?
- dilyevsky 5y ago1. Leadership prefers to blame cloud when things break rather than take responsibility. 2. Cost is not an issue (until it is but you’re already locked in so oh well) 3. Faang has drained the talent pool of people who know how
- marcosdumay 5y agoIf you trust them at this point, you have not being paying attention, and will probably continue to trust after this.
- jeffrallen 5y agoWell, no big deal, there's not really a lot of trust there to destroy...
- dekhn 5y agoUhhhhh... what if the monitoring said it was hard down? They'd still not show red?
- choeger 5y agoProbably they cannot. They outsourced this dashboard and it runs on AWS now ;).