5 ms·
Can anyone explain why status pages are so difficult. Theres even statups like status.io dedicated to this one thing. It really does seem that anytime there is
by s_dev 6y ago
Can anyone explain why status pages are so difficult. Theres even statups like status.io dedicated to this one thing.
It really does seem that anytime there is an outage more often than not the status page is showing all green traffic lights. Making it redundant as a tool to corroborate whats happening.
How did AWS status page compare with status.io/aws?
- xyzzy123 6y agoWhen your company gets sufficiently large, outages become political. Failure happens at the speed of computing but agreeing that something is failing in a way that customers need to be told about is a slower process. Even when status pages are fully automatic (rather than manually updated), there will tend to be gaming of the metrics that constitute that. Ideally you would just be monitoring your SLOs and publishing that to customers... that doesn't seem to be how it works, anywhere.
- deleted 6y ago[deleted]
- freehunter 6y agoAnd not just outages, but security incidents. I’ve worked at/with/for many companies as both an employee and a consultant where the top priority wasn’t to have fewer security incidents, but to have fewer security incidents that would require disclosure. Publicly disclosing an incident to a customer is embarrassing and potentially damaging but almost equally as damaging is telling other teams you had an incident. Now anything that goes wrong is your fault by default because “it’s probably related to that incident” and any new security policies are blamed on the other team: “we wouldn’t have to do that if Ops didn’t mess up last month”. The answer to “is this service suffering an outage” is seriously complex and hard to determine. The answer to “is this a security incident” is 10x harder and 100x more political because the industry is still just so wildly immature.
- drchopchop 6y agoAdditionally, you're penalized for doing it "right", because you're often competing against companies which rarely say that anything's wrong (ahem, Mailchimp). You look worse, because you're being transparent about service status, which creates the perception that you're generally less stable.
- dexterdog 6y agoAll of those are reasons that the determination of status should be totally independent of the company technically and legally.
- PragmaticPulp 6y agoMany companies tie uptime and outages to performance reviews, either directly or indirectly. Admitting that your services are down could be costly to your career progression and bonus. When people know this, they go to great lengths to avoid admitting fault. Updating the status page is the first admission of fault. The longer the status page shows an outage, the worse it gets. I worked with an ex-Amazon engineer at a previous company. After each outage, he would spend days or weeks writing long reports explaining how the outage was not his fault. He didn't care so much about downtime so much as not getting blamed for outages. Predictably, this was terrible for team morale and most of his team members ended up quitting. If anyone else finds themselves in this position, the solution is have another team responsible for monitoring uptime, and to rate teams on how quickly they acknowledge outages. Once the response time and accuracy of your status page becomes a performance metric, people are less likely to play games with it.
- opmac 6y agoIt is kind of perplexing that AWS dogfoods its own status page. I remember during the massive S3 outage a few years ago that their status page remained green almost the entire time because the red/green/blue icons for the status was stored in... wait for it... S3. You'd think they would have learned from that.
- Twirrim 6y agoThey did. It came up in the post incident report, and senior leadership kicked off work to have it run on its own distinct infrastructure so that this wouldn't happen again. If you look at where the content on https://status.aws.amazon.com/ https://status.aws.amazon.com/ is actually hosted from you'll see things like the status icons are all hosted under the same domain, e.g. https://status.aws.amazon.com/images/status1.gif https://status.aws.amazon.com/images/status1.gif https://status.aws.amazon.com/images/status0.gif https://status.aws.amazon.com/images/status0.gif etc. If you look at the source code for the site, you'll again see that everything is hosted from the same domain. One of their main goals was to ensure that it could never go wrong that way again.
- opmac 6y agoK so they avoided that problem, but something similar has obviously gone wrong again, considering that Kinesis had been partially or fully down for almost an hour before the status page got their first update. And the fact remains that currently an outage of AWS's own infrastructure is impacting AWS's ability to status updates on its own status dashboard. It's just seems so... amateurish.
- Twirrim 6y agoThat's incredibly annoying, given the mandate the replacement service had. I'd be curious to be a fly on the wall during the next Ops meeting when it comes up that yet again the status dashboard got made in a way that makes it hard to update during an outage.
- 6y ago
- bak3y 6y agoAny time companies have SLA's where money is on the line if they admit they're having an outage, they're going to be delayed on updating a status page.
- sailfast 6y agoIt says on this outage page (as of 11:11 ET) that the problem with Kinesis is also causing problems updating this outage dashboard which may explain the delay?
- latch 6y agoStatus pages, like SLAs, are sales tool - not engineering tools. At best, they are there to help decision makers go through their checklist. At worse, they exist to deceive.
- jmartens 6y ago1 million percent! Which makes me wonder, why do we all rely on status pages rather than solve the problem ourselves in ways that don't require us to rely on the vendor?
- Twirrim 6y ago> Can anyone explain why status pages are so difficult. What is an outage? When does an outage reach sufficient scale that updating the status page is the right thing to do? I used to work for AWS, and now work for another cloud provider. One thing that's hard to communicate is the sheer scale that these services operate at, what that means architecturally, and how they tend to break. Outages, even just slight degradation, occurring on a whole service scale are very rare. I would argue from my experiences there that most incidents affect less than 10% of any given service's customers. Whether it gets noticed in part depends on who is encompassed by that percentage. What is very often the case is that a subset of customers get impacted to some degree during any given incident. That can be even things like single percentage of customers or less, but be an incident that has all hands to deck and the entire management chain of the service aware and involved in. At what percentage do you draw the line and say "Yes we need this many percentage of our customers to be affected before we post a green-i" (AWS terminology for the first stage of failure notification). How do you communicate that effectively to customers, in such a way that doesn't suggest your service is unreliable when it really isn't. The moment you post a green-i or above, customers start blaming you and your service for problems with their infrastructure that are not caused by it. If you're looking to use a service and go look at the status history and see it filled with green-i or similar, are you likely to trust it? No. Even if those green-i's were for impacts on a limited subset of customers. AWS wrestled with this a bunch about 5-6 years ago. There were no end of discussions during the weekly ops meetings with senior leadership, directors and engineers across the company. Everyone wants to do the right thing and make sure customers get an accurate picture about the health of the service, without giving the wrong impression. In the end they opted to move towards having personal notifications for outages, and build tooling to help services quickly identify which customers are being affected by any particular incident and provide personalised status pages for them that can be way more accurate than any generalised status page.
- Kinrany 6y agoPosting percentages instead of green/red would fix all of these, no?
- 6y ago
- darkcha0s 6y agoI completely agree, but can we talk for a second how absurd it is charging 90$ for essentially a service that just pings your infrastructure?
- acdha 6y agoTry undercutting it. At some point you’ll learn that the problem isn’t that simple, operations is a key part of the product and isn’t free, and people expect support for important services.
- jmartens 6y agoExcept, the option to ping a service in order to programmatically inform a status page is almost never used. The dirty secret of status pages is that they are almost always manually updated, typically only when a very high bar is met, and after senior managers, sometimes even comms people, approve it.
- bird_monster 6y agoThey aren't difficult. Amazon has no interest in having a working status page. Amazon would prefer the appearance of always green checkmarks over actually having a status page.
- Griffinsauce 6y agoThis is why my instinct is to check Twitter feeds of the related service first. So far in several years of experience it has been more informative and helpful than a status has ever been. It's a sad state.
- jessaustin 6y agoOne never thought we'd see the day... Twitter, that storied home of the whales of fail, is the reliable service.
- swyx 6y agothere's also https://stop.lying.cloud/ https://stop.lying.cloud/
- cheeze 6y agoIronically I can't even load that page
- Thaxll 6y agoIt's not that easy to quantify how down a service is at the scale of AWS, for example Cognito has issues, does it means every services that rely on that have issues, what is the impact etc...