5 ms·
I wonder why the status just doesn't ping github.com for 200. That seems easy to do.
by ergocoder 2y ago
I wonder why the status just doesn't ping github.com for 200. That seems easy to do.
- tinyhitman 2y agodelaying SLA
- sebmellen 2y agoThis is at least a multi-million dollar payout (if they admit to it). All GitHub Pages say > We're having a really bad day. > The Unicorns have taken over. We're doing our best to get them under control and get GitHub back up and running.
- cbates 2y agoSeems slightly unproffesional for a massive company like Github/Microsoft.
- xp84 2y agoI disagree. This hurts no one, and not everything needs to be sanitized and painted over with bland corporatespeak.
- majewsky 2y agoI don't think they were asking for corporate speak. But at least I would find a plain technical error message like "cannot contact file server" much more respectable than something like "unicorns are hugging our servers uwu".
- COMMENT___ 2y agoThis “ironic” and “humorous” style of errors and UI captions is the actual new corporate speak. I’d prefer dumb error messages rather than some shit someone over the ocean thinks is smart and humorous. And it’s not funny at all when it’s a global outage impacting my business and my $$$.
- colimbarna 2y agoIt's closer to the truth than you usually get. They're having a bad day, it's completely true. It's the start of my day, but I guess this is the middle of the night for them. There's no such thing as unicorns, but that just highlights the metaphorical nature of the remaining claim - getting Unicorns under control means solving their problems. Normally "professional" corporate speak means avoiding saying anything whose meaning is plain on its face and disconfirmable while avoiding the implication that the company is run and operated by humans. This is a model. (Obviously the came up with the message in advance, which just goes to show that someone in the company is well enough rounded to know that if it is displayed, they're having a bad day.)
- wrs 2y agoGitHub is (was?) a Rails application, so it was probably originally running behind Unicorn [0], if it isn’t still. So the unicorns are (were) real. [0] https://en.wikipedia.org/wiki/Unicorn_(web_server) https://en.wikipedia.org/wiki/Unicorn_(web_server)
- ljahier 2y agoAt the moment, all github services seem to be restored, and the github status indicates that the problem is still ongoing. I don't think it's related to the SLA, but rather to the monitoring, which is not live. There are a few minutes of delay.
- fragmede 2y agofrom where? they don't only have one load balancer, so you'd still have the problem of the page showing green when it's not loading for some folk?
- ergocoder 2y agoAt Github's scale, why wouldn't they put a ping monitor from every continent at least? Then, you would show the status based on the continent.
- intelVISA 2y agoThat would be self-defeating given that it's a Rails app.
- fragmede 2y agoWhere on the continent? GitHub is undoubtedly doing blackbox testing internally and has multiple such monitors but that's not going to capture every customer's route to them, leading to the same problem - customers experience GitHub being down, despite monitoring saying it's mostly up. Thus the impass. Even doing whitebox testing, where you know the internals and can this place sensors intelligently, even just for ingress, you're still at the mercy of the Internet. If a sensor that's basically in the same datacenter says you're up, but the route into the datacenter is down, then what? multiply this by the complexity of the whole site, and monitoring it all with 100% fidelity is impossible. Not that it's not worth it to try, there's a team at GitHub that works on monitoring, but beyond motivation about keeping the SLA up, as a customer, unless you notice it's down, is it really down? In a globally distributed system, downtime, except for catastrophic downtime like this, is hard to define on a whole-site basis for all customers.
- laserlight 2y ago> 100% fidelity is impossible I don't think anybody asked for 100% fidelity. We are talking about a complete outage that affected at least North America and Europe. If the status page shows green in such a case, its fidelity is around 50%. People expect better from GitHub.
- bigiain 2y agoTo be fair - I really couldn't care less is the homepage is loading or not. So long as I can fetch/commit to my repos, pretty much everything else is of secondary, tertiary, or no real importance to me. (At work, I do indeed have systems running that monitor 200 statuses from client project homepages, almost all of which show better that 99.999% uptimes. And are practically useless. Most of them also monitor "canary" API requests which I strive to keep at 99.99% but don't always manage to achieve 99.9% - which is the very best and most expensive SLA we'll commit to.)