4 ms·
GitHub used to have a pretty cool status page, with all kinds of real time graphs. Does anyone know what happened to it? Since it makes me really sad that this
by rollulus 6y ago
GitHub used to have a pretty cool status page, with all kinds of real time graphs. Does anyone know what happened to it? Since it makes me really sad that this status page is a plain lie, I had to visit HN to get the confirmation that they are having issues again, and that it just wasn't only me.
- uberman 6y agoThe status page clearly states (when I look) that: Incident on 2020-07-15 15:41 UTC We are investigating reports of degraded performance. Posted 9 minutes ago. Jul 15, 2020 - 15:41 UTC
- rollulus 6y agoCurrently it does, indeed. From my Slack logs, at 15:00 UTC I noticed problems. I'm pretty sure that message is manually created, at least 41 minutes after the fact.
- luckylion 6y agoThat's the most annoying thing. Usually when I get notifications from monitoring about some issue, the first thing I do is check the vendor or provider's status page to see whether it's an issue on their end. If there's nothing, I go and investigate. Recently, more and more of them take 10-15 minutes until they mention a service outage. I don't work in super HA, I don't want to get an alarm because a single ping failed etc, so I'm lenient and have a few minutes of a delay in alarms. If I'm writing an internal incident report before the official status page is updated, that's bad. This seems similar: external users noticing the outage and posting on HN before GitHub notices & acknowledges it.
- deathanatos 6y agoWe have the GitHub status RSS integrated into our Slack channel. One of my company's engineers noticed the outage at 15:06 UTC; the RSS feed picked it up at 15:49 UTC, though the message text says it was from 15:41 UTC. (And I think RSS polls, so there's some inherent lag, so I'd take the 15:41 UTC timestamp.) The half hour in between was us debugging, thinking it was us.
- hinkley 6y agoThe last straw that got me out of mobile was working at a place with bad engineering discipline (or more precisely, bad management of engineering discipline). They were either paranoid or just didn't trust the team, and every time there was a blip in traffic someone in management would ride the engineers until they could prove it was on the other end. It almost always was. When I later saw the "stack trace or GTFO" comic I had a pretty clear idea what the author was feeling. Eventually they rearranged the cube walls so management had to get more exercise to come harass the team. Yes, it was better use of space and the windows (in part due to my input), but that's not why the 2 people who started disassembling the cubes were doing it. "Fit of pique" is a phrase I don't get to use as often as I like, but that's what it was, whether cooler heads legitimized it or not. Oddly, someone tried to blame my failure to convert to FTE on my interactions with one of those two engineers. He was all bark and not that much bite though. I could already handle him almost as well as anybody else and I was the new guy. No, they were trying to get everyone pagers and if that same kind of interaction happened at 2 am, I was gonna say something that got me fired. Found a much better offer and I stayed at the next place for 5 years, working on a surprising array of things and nobody ever said the p-word to me since.
- originof 6y agoI read an article about it, look for "status page evolution" https://nimbleindustries.io/2020/06/04/has-github-been-down-more-since-its-acquisition-by-microsoft/ https://nimbleindustries.io/2020/06/04/has-github-been-down-...
- kohtatsu 6y agoFrom the first two graphs it looks like they are a lot less liberal about using "down" instead of "warn".
- hinkley 6y agoThe best triage policies I've ever gotten to work with had severity and priority separated. Severity went something like this (sometimes the numbers flip which always confuses at least 20% of the team about whether things are almost normal or people are hunting each other for sport). 1: data loss 2: some workflows blocked 3: some workflows unavailable w/ workarounds (ie other routes) 4: Everything else except 5: Irritations Having a UI break but the underlying functionality is still working is not good but people can still do their jobs, if more slowly. It's important to classify these separate from S2 and S4. There is urgency but don't panic. Go eat lunch or have your planning meeting, then go fix it. If data is getting lost ain't nobody doing nothin' until we figure it out, and then some people can go back to work but don't interrupt the people still working on it. I think the problem is that so many metrically dysfunctional people, to the point of cliché, have rationalized that an S2 means that only 20% of our customers can't do their jobs so we are degraded but still working normally, when really a yellow status should be at S3, while S2 should be at least orange although those affected will be upset that it's not red. Over time that 20% will shift around to most of your customers. Eventually several times, and then you'll wonder why everyone is talking trash about you on HN. It's not like that many people were affected!
- hinkley 6y ago> But that could be all a part of coordinated effort to be more transparent about their service status, an effort that should be applauded. Microsoft could be pushing for transparency. Or people are more relaxed about transparency now that GitHub has its exit. How long did GitHub know they were looking to be acquired? Maybe this analysis should look at a longer time interval..