3 ms·
> There is a non-zero chance that something bad could happen with the domain (...) If something happens with the domain that prevents you from accessing the se
by rewmie 3y ago
> There is a non-zero chance that something bad could happen with the domain (...)
If something happens with the domain that prevents you from accessing the service, the page works as intended by being down. I mean, what do you think users will say when they go to the domain and their browser shows an error?
- viraptor 3y ago> what do you think users will say when they go to the domain and their browser shows an error? - Is my internet down? Can I access any page? - Is that a region accessibility issue for my location? - So they killed the status service and I won't know when the main service is back. The outage scope and estimated time to be back online may not matter much if you just want to check the latest memes. But if you're coordinating some product launch / promotion, this may be very important info. Even more so if that idea of "everything app" ever materialises and the question becomes "so when can I make that payment".
- rewmie 3y ago> - Is my internet down? Can I access any page? ...and in a few seconds the same user opens google.com. What then? > - Is that a region accessibility issue for my location? Why does a user care? A user only cares if they can access the page from where they are at that moment. > - So they killed the status service and I won't know when the main service is back. Users check if the main service is back by hitting F5 to try to access the main service. Either it's up or down. If they can't access the service, it makes no difference if they can access the status page or not, because they already have their answer. I think you're trying very hard to overthink and overcomplicate something that's quite simple and straight-forward. All a canary does is check if the service is up, and possibly track uptime as well.
- willsmith72 3y ago> Either it's up or down. If they can't access the service, it makes no difference if they can access the status page or not, because they already have their answer. This is just categorically false. The difference between "something went wrong", "we're looking into it", "we've found the problem and are working on a fix", and "we have a fix and are deploying it region-by-region" is huge. If an upstream dependency of mine is having issues, I need to communicate it to my clients/users as well, and they also want to know this stuff. E.g. a lot of people use Slack with external users for support and collaboration. The scope of the outage will determine if I should try to onboard them to a new system as quick as possible or wait it out. But this isn't just about Slack. It could be any service operating at any scale, with varying levels of support. Employees might be asleep, it might be that just my subdomain/project is having issues, it could be planned maintenance, maybe an issue with the hosting provider. As a customer I want to know all of that, and as a business I feel obliged to provide it.
- rewmie 3y ago> The difference between "something went wrong", "we're looking into it", "we've found the problem and are working on a fix", and "we have a fix and are deploying it region-by-region" is huge. It's also a figment of your imagination. No service provider provides those fine-granular updates. Even AWS shows their services as up when they are clearly major outages going. The only reliable services tracking outages are third-party canary services that rely on crowdsourced reports, and obviously those don't come even close to providing any insight onto progress. I think you're confusing your own wishful thinking with real-world implementations of status reporting services.
- willsmith72 3y agoI think you haven't worked with or on real production systems, or with real paying clients
- viraptor 3y ago> No service provider provides those fine-granular updates. I wrote some of those, so... yes they do. You just didn't get to experience them.
- willsmith72 3y agoThat's not at all what a good status page is. If I'm a business using a service, and I'm having issues with it, I need to know: - If they know about it - If they have a fix ETA - Any other problems related to their service which I haven't noticed And many more things, which actually need a status page not an "error"
- rewmie 3y ago> If they know about it What exactly leads you to believe that each and every employee of Slack would fail to notice that each and every single service accessed through the domain slack.com was down? > - If they have a fix ETA I never saw a single status page that provided that. At all. In fact, the most reliable status monitoring services tend to be crowdsourced third-party services, which obviously do not track ETAs for fixes. > Any other problems related to their service which I haven't noticed I see no reason why that could not be provided through the company's domain.
- HPsquared 3y agoDomain / IP could perhaps be inaccessible only from certain locations, and staff could be unaware in that case.
- rewmie 3y ago> Domain / IP could perhaps be inaccessible only from certain locations, and staff could be unaware in that case. Your hypothetical scenario either involves a regional deployment being down, which I would be very surprise if any non-amateur setup didn't tracked with canaries, or if there's a problem with how the internet is being routed, which is something that's beyond the control of any team but canaries would also flag.