4 ms·
Even better IMO is this status page: https://mrshu.github.io/github-statuses/ https://mrshu.github.io/github-statuses/ "The Missing GitHub Status Page" with ov
by mholt 6mo ago
Even better IMO is this status page: https://mrshu.github.io/github-statuses/ https://mrshu.github.io/github-statuses/
"The Missing GitHub Status Page" with overall aggregate percentages. Currently at 90.84% over the last 90 days. It was at 90.00% a couple days ago.
- skipants 6mo agoThese are two pages telling two different things, albeit with the same stats. The information is presented by OP in a way to show the results of the Microsoft acquisition.
- montroser 6mo agoIt has been pretty rough. Their own numbers report just a single `9` for Actions in Feb 2026 with 98% uptime. But that said -- I don't get the 90% number. Anecdotally, it seems believable that 1 in 50 times (2%) in Feb that Actions barfed. Which is not very nice, but it wasn't at 1 in 10 times (10%).
- verdverm 6mo agoIt looks like the aggregate stats are more of a venn diagram than an average. So if 1/N services are down, the aggregate is considered down. I don't think this is an accurate way to calculate this. It should be weighted or in some way show partial outages. This belief is derived from the Google SRE book, in particular chapters 3 (embracing risk) and 4 (service level objectives) https://sre.google/sre-book/embracing-risk/ https://sre.google/sre-book/embracing-risk/ https://sre.google/sre-book/service-level-objectives/ https://sre.google/sre-book/service-level-objectives/
- mort96 6mo agoI mean I think it's useful. It answers the question, "what percentage of the time can I rely on every part of GitHub to work correctly?". The answer seems to be roughly 90% of the time.
- naniwaduni 6mo agoNobody cares about every part of GitHub working correctly. I mean, ok, their SREs are supposed to, but tabling the question of whether that's true: if tomorrow they announced a distributed no-op service with 100% downtime, you should not have the intuition that the overall availability of the platform is now worse.
- verdverm 6mo agoI don't use half of the services, the answer is not straight forward https://mrshu.github.io/github-statuses/ https://mrshu.github.io/github-statuses/
- ablob 6mo agoIf you're using all services, then any partial outage is essentially a full outage. Of course, you can massage the numbers to make it look nicer in the way you described but the conservative approach is better for the customers. If you insist, one could create this metric for selected services only to "better reflect users". That being said, even when looking at the split uptimes, you'd have to do a very skewed weighting to achieve a number with more than one 9.
- verdverm 6mo ago> That being said, even when looking at the split uptimes, you'd have to do a very skewed weighting to achieve a number with more than one 9. It's definitely bad no matter how it you slice the pie. If GH pages is not serving content, my work is not blocked. (I don't use GH pages for anything personally)
- marcosdumay 6mo agoThat's how you count uptime. You system is not up if it keeps failing when the user does some thing. The problem here is the specification of what the system is. It's a bit unfair to call GH a single service, but it's how Microsoft sells it.
- verdverm 6mo ago> That's how you count uptime. It's not how I and many others calculate uptime. There is not uniformity, especially when you look at contracts.
- tbossanova 6mo agoAs a “customer”, I consider github down if I can’t push, but not down if I can’t update my profile photo (literally did this today, sending out my github to potential employers for the first time in a long time). This stuff is notoriously hard to define
- formerly_proven 6mo agoIn a nutshell, why would the consumer care (for the SLO) care about how the vendor sliced the solution into microservices?
- verdverm 6mo agoIt will depend on the contract. When I was at IBM, they didn't meet their SLOs for Watson and customers got a refund for that portion of their spend
- bandrami 6mo agoThinking back to when I was hosting, I think telling a customer "your web server was running fine it's just that the database was down" would not have been received well.
- fontain 6mo agoAn aggregate number like that doesn’t seem to be a reasonable measure. Should OpenAI models being unavailable in CoPilot because OpenAI has an outage be considered GitHub “downtime”?
- fwip 6mo agoI think reasonable people can disagree on this. From the point of view of an individual developer, it may be "fraction of tasks affected by downtime" - which would lie between the average and the aggregate, as many tasks use multiple (but not all) features. But if you take the point of view of a customer, it might not matter as much 'which' part is broken. To use a bad analogy, if my car is in the shop 10% of the time, it's not much comfort if each individual component is only broken 0.1% of the time.
- remus 6mo ago> But if you take the point of view of a customer, it might not matter as much 'which' part is broken. To use a bad analogy, if my car is in the shop 10% of the time, it's not much comfort if each individual component is only broken 0.1% of the time. Not to go too out of my way to defend GH's uptime because it's obviously pretty patchy, but I think this is a bad analogy. Most customers won't have a hard reliability on every user-facing gh feature. Or to put it another way there's only going to be a tiny fraction of users who actually experienced something like the 90% uptime reported by the site. Most people are in practice are probably experienceing something like 97-98%.
- fwip 6mo agoSorry, by 'customer' I meant to say something like a large corporate customer - you're buying the whole package, and across your org, you're likely to be a little affected by even minor outages of niche services. But yeah, totally agree that at the individual level, the observed reliability is between 90% and 99%, and probably toward the upper end of that range.
- deleted 6mo ago
- goodmythical 6mo agoholy shit that's nearly five weeks of down time. Well, I mean, I guess that's fair really. How long has github been around? Surely it's got five weeks of paid time off by now...