6 ms·
I was seeing PRs failing to update & webhooks failing to trigger upon pushing code for 30 minutes before GH's status page acknowledged anything. I'm surprised t
by arilotter 6y ago
I was seeing PRs failing to update & webhooks failing to trigger upon pushing code for 30 minutes before GH's status page acknowledged anything. I'm surprised they don't have monitoring in place that would catch webhooks failing within minutes of the failure beginning.
- malux85 6y agoAt large bureaucratic organisations there's often political implications of changing the official status, so often it lags behind reality until it cannot be swept under the rug anymore Not saying it's right, just an observation
- KingOfCoders 6y agoAs a CTO responsible for an often failing eCommerce website that lost millions when down and which I took over, I fell in the same trap of trying to sweep things under the rug. Until I decided to no longer do that and my life improved considerably.
- malux85 6y agoYeah it's very frustrating - especially if your customers are technical, they are seeing the errors and the status page says everything is fine. I've seen status pages and error counts tied to bonuses, which only caused a giant mess of bad incentive alignment and internal lies, customers are unhappy, developers are unhappy, management are lying to upper management, it's so much easier to focus efforts on real problems and just be honest and improve. Thank goodness I dont work there anymore (cough cough Google)
- penagwin 6y agoDo these companies not have live error reporting and tracing? Like surely github got alerts that things weren't working? Why don't they just hookup their status page and their alerts? Or is it a political/relationship thing, and they want to have a human give out the status page updates? This could have been caught with a cron job and some curl requests :\
- mjayhn 6y agoIn all honesty they're typically just banking on people not noticing it and are trying to make it as little of a fuss as possible and get it up before it gets to twitter. The problem is when it's not just a small blip and they haven't addressed it and it goes mainstream and is still down, it just leads to concerns about transparency. Building infra I have to work around all sorts of 3rd party services going out or having blips throughout the day, docker registries, caches, bgp, etc., it's totally an expected part of infra design but not every team has the time or need to build in the resiliency. I see tons of outages that never get reported or IMO aren't reported adequately enough. With that said, I'm no angel, I get all my service down notifications through slack, so when slacks down..
- JMTQp8lwXL 6y agoWhat's the point in having a status page if it's a political artifact? In that case, it serves zero customer value.
- opportune 6y agoIt’s still an indicator that “no you’re not crazy, we’re having issues on our end” but not a foolproof one. It’s kind of sad that twitter is usually the best place to confirm an outage as it begins, rather than the software providers themselves. I assume if they actually exposed global availability metrics in most cases it would not look as good as they would want it to
- paulie_a 6y agoDown detector is based on twitter complaints and is pretty damn accurate.
- tfolbrecht 6y ago3 words: Service Level Agreement I've caught a big cloud provider not reporting a degraded service, I assume they knew but politics and $ come in and it's easier to just gaslight everyone. I get it, but my frustration is worth loosing a trailing 9. I think there should be some 3rd party continuously testing APIs. Degraded states are downtime!
- tdeck 6y agoHonestly I think this can be true at any size organization. Small startups often take the approach of "let's hope no one noticed while we try to fix it", it's just that they have fewer users to notice so it's more likely to work.
- chrispauley 6y agoHad the same issue this morning. The lagging status always causes the issue of "is it you, me or GitHub?" snaffoos. Really annoying to have these issues so consistently. Would switch to gitea or similar in a moment given the choice.
- hyperdimension 6y agoJust a friendly correction: `SNAFU:' Situation Normal: All Fucked Up
- hinkley 6y agoThis misspelling brought to you by foobar. Foobar: for when you are too polite to say FUBAR (Fucked Up Beyond All Recognition).
- grecy 6y agoMy two favourites lakes in the Yukon - SNAFU and TARFU (Things are Really Fued Up). Named, of course, by the Army when they built the Alaska Highway.
- chrispauley 6y agoI had never given any thought to this word as I had heard and used it since childhood, had no idea that was the origin!
- deleted 6y ago[deleted]
- njsubedi 6y agoThey used to have real-time graphs and stuff on their status page. That was a thing of the past; with more distributed system, they're probably not sure if the service is down everywhere. If a node somewhere is still up, they might consider the service up. I don't know much, but it's up to the kind of downtime measurement system they have.
- emilfihlman 6y agoThat's the way it is _everywhere_. For example status.digitalocean.com is _not_ real time, it's manually updated. And it's irritating as fuck.