6 ms·
My action failed with "Unexpected error fetching GitHub release for tag refs/heads/master: HttpError: Sorry. Your account was suspended" Which certainly made m
by a10c 4mo ago
My action failed with "Unexpected error fetching GitHub release for tag refs/heads/master: HttpError: Sorry. Your account was suspended"
Which certainly made me shit myself, briefly.
- grim_io 4mo agoA brownout redefined.
- lachieh 4mo agoGood thing I'm wearing my brown pants today.
- DonHopkins 4mo agoShitHub https://www.youtube.com/watch?v=LGeOee7x5lY https://www.youtube.com/watch?v=LGeOee7x5lY
- drcongo 4mo agoSame. It's weird how I always find out that GitHub is down before GitHub does. Took 15 minutes before it appeared on githubstatus.com
- jaapz 4mo agoAll these monitoring rules are of the format "when 500 errors > baseline for x minutes". Otherwise you'd have monitoring alerts every second. So it is normal for users to already see errors before github officially counts it as an outage.
- echelon 4mo agoIn a high performance service with good maintenance and upkeep, you page for all 500s. A noisy pager forces the team to fix the 500s. Maybe the Github Actions infrastructure isn't run like that. edit: my oncall rotation notified on all 500s, 24/7, not just rates - https://news.ycombinator.com/item?id=48279262 https://news.ycombinator.com/item?id=48279262
- TheDong 4mo agoDo you know of a single service at a single company that actually does that? I know all of Gmail, every GCE service I can think of, every AWS service I can think of, Amazon.com, Netflix, and Github all do not page on just a single 500. I know none of those are particularly "high performance" though. Curious where your experience is coming from.
- CBLT 4mo agoI've been oncall for a different G service that nearly paged on every error. It used the standard error budget tooling, but on hundreds of user buckets because the engineering around locality-specific configuration was... suspect. Many of these buckets had single-digits user. If a user was on their phone and lost signal, I was paged. Very poor oncall experience.
- echelon 4mo agoI worked at a large fintech moving billions of dollars in volume a day. I had a fairly long tenure, where I maintained multiple key services in critical online payments flow. Authentication, authorization, core business and risk data, as well as some cross-cutting control plane stuff, etc. You needed one or more of our services to take a payment, serve any request from the employee dashboard - pretty much everything hit our services. The entire company ground to a halt without my team. We paged for every single 500. In instances where a particular class of 500 was spurious or not worth fixing, we would leave it acked or mark it as noise. But typically we'd just put in a fix as soon as possible so we didn't page. Our graceful shutdown and traffic shaping stack was great, but occasionally we'd get a few pages during deploys or failovers. Oncall was typically not bad, but when it did get bad it was terrible. I've been involved in huge outages that cost hundreds of millions of dollars. Usually it was the fault of multiple teams having compounding runaway failures rather than one service or bug in particular. It's inexcusable to have a customer's payments not go through. We engineered around resilience. We had strict five nines SLAs and p99 targets and evaluated our adherence with even the smallest partial outage. Hundreds of other services depended on ours, and downstream impacts were huge, so we had to keep a tight ship. We didn't have "business hours"-only paging either as our platform was available globally, including a heavy install base in Asia.
- hnlmorg 4mo agoYou'd expect them to be monitoring more than just the HTTP response codes from user requests for precisely this reason. If the first they hear of an outage is when user requests start to fail, then that's a failure in their monitoring as well. But effective monitoring is harder than people assume.
- dncornholio 4mo ago> If the first they hear of an outage is when user requests start to fail, then that's a failure in their monitoring as well. Isn't that what monitoring actually is? The issue seems to be in their testing, not monitoring.
- hnlmorg 4mo agoNo, monitoring for HTTP response code is a subset of observability and not one that generally gives you the best insights into which subsystems are misbehaving nor why. There are synthetic tests, where you can generate API request calls or even simulate an entire user journey. These allow you to control the user agent, the payloads, and thus you know anything errors back are actual errors. These are triggered by the observability platform (think like running a cron-job) and thus you're not tied to user activity to see when problems arise. There are other metrics outside of HTTP response codes too. Think like free RAM, CPU usage, disk space, etc. This is just naming some obvious ones because these types of metrics are generally bespoke to the type of application your monitoring. And with these types of monitors, you'd not just have an alert when things have failed, but ideally have alerts when an irregular trend is showing that things are likely to fail too. This latter type of monitors helps you get ahead of the problem before it become customer facing. Then you have more traditional stuff like logs. This will also be bespoke to the application. But you'd expect errors in logs to get surfaced quickly. Assuming Github have good hygiene in what's being logged. Tie that up with APMs, RUM, and other goodies like that and you'll have diagnostics to investigate issues when they appear. (this is just a super high level view of observability too)
- lokar 4mo ago
- logifail 4mo ago> All these monitoring rules are of the format "when 500 errors > baseline for x minutes". Otherwise you'd have monitoring alerts every second. So it is normal for users to already see errors before github officially counts it as an outage. Is it true that official service status pages are updated automatically?
- baby_souffle 4mo ago> it true that official service status pages are updated automatically? Depends. Typically no because there’s an art to crafting the actual message around impact… but sometimes yes it is automated
- logifail 4mo ago> Typically no because there’s an art to crafting the actual message around impact I was thinking more of needing to notify/get sign-off from management...
- baby_souffle 4mo ago> I was thinking more of needing to notify/get sign-off from management... Yeah, that's usually part of it. Precise language matters a TON when you might have some expensive breach-of-SLA terms. Sometimes the people first responding don't even have the full picture yet and can't fully articulate the impact so they leave it vague.
- registeredcorn 4mo agoI'm not arguing with what you're saying, but it does make me wonder: What exactly is the point of the status page, if "it is normal for users to already see errors before GitHub officially counts it as an outage"? Is it more so to have something to link to for managers who aren't using the service have a pretty bar to look at and feel like they are "doing something"? Or is it more of a kind of a way to prevent confirming what you already suspect to be true. E.g. "Huh. Me and Jim are seeing problems. How about you Tom? Oh wait, crud. The service page is confirming it's down now. Never mind! Who wants coffee?!"
- filleduchaos 4mo agoThere is oddly enough a middle ground between "zero errors whatsoever" and "outage".
- deleted 4mo ago[deleted]
- simonjgreen 4mo agoMore likely that 'update the Status site' lives a long way down their incident response plan, and they have alarms going off well before that
- jordemort 4mo agoyeah I mean a company the size of GitHub certainly can’t be expected to have enough staff to walk and chew gum at the same time
- swiftcoder 4mo agoIf it's like other BigTechs I have worked at, you need director-level signoff and comms team approval to post an outage notice
- PunchyHamster 4mo agoit should be automatic tho. Probably isn't so they can at least get the one nine on availability
- simonjgreen 4mo agoMarketing definitely takes interest in status sites
- re-thc 4mo ago> It's weird how I always find out that GitHub is down before GitHub does No, it's not. Official updates = potential SLA penalties. Always requires approval.
- drcongo 4mo agoThis is the most plausible reply.
- deleted 4mo ago[deleted]
- chrisjj 4mo ago> githubstatus.com There's a threshold. It shows only once 1000 users complain. /i
- deleted 4mo ago[deleted]
- dvduval 4mo agoYes, Thais can be be really frustrating when you’re trying to get work done. There needs to be more competition and better alternatives and the LLMs need to offer easier connection to these alternatives.
- weird-eye-issue 4mo agoWhat do the Thai people have to do with this? :(
- superxpro12 4mo agoReminded me of the "Thai Fighter" joke from family guy's star wars spoof lol
- denisw 4mo agoPretty sure that they wanted to write "this", typed something different by accident, and auto-correct struck.
- weird-eye-issue 4mo agoOh gee thanks
- deleted 4mo ago[deleted]
- neya 4mo agoIt's an eye opener. Think about it - today, it was a mistake. But, what if it really happened? What if you really lost access to all your years of hard work? It's a wake up call. A blessing in disguise to store what matters to you the most locally, backed up offline. Never trust any single provider. Be it MS or Google or Apple. RAID is the way.
- onion2k 4mo agoPeople should use something that keeps a local copy of their code and just copies it to Github and to other contributors with a sync process to push and pull changes. Some sort of 'distributed source control system' maybe. Then people would only need a 'hub' to connect to people, and it'd be easier to move somewhere else.
- coldpie 4mo agoThis gets tiresome. Github is a lot more than a host for Git repositories. If you want to suggest that people use something else, you need to suggest a replacement that has the features people use Github for.
- ornornor 4mo agoIncreasingly less and less so as they “upgrade” their offering and have more and more downtime.
- doctorpangloss 4mo agoyeah, #1, it is free private file storage, and #2, it's a download portal for free as in beer software replacing paid offerings. that's what it is for 99.99% of people. being a host for git repositories has never been its core competency. neither has its groupware offering. does it even serve OSS well? a very interesting criteria is, "Have mature or adopted end-user-facing OSS recently merged a large PR from an unallied contributor?" The answer is overwhelming no. This is why there is so much innovation in this space.
- danudey 4mo ago
- ridiculous_leke 4mo ago> Which certainly made me shit myself, briefly. Can you sue companies for inducing such anxiety?
- Imustaskforhelp 4mo agoIANAL, but I can probably imagine a case being made if a person really got so stressed that for example any health condition got invoked from the stress. It might be up to the lawyer to explain how exactly the service caused the stress and its direct relation to health condition though and up to the judge. but I suppose that there might be some terms of conditions within using github (ahem Microsoft) that you can probably not sue them for something like this. It really depends upon the severity of situation (imo) For example, if a person had any heart condition and they got so stressed because of an error at github (which to be fair, I can understand the stress part, imagine losing some part of your software because it was on github and the amount of direct damage to livelihood if your income depended on it) and I think that the judge might have to be in just the right technical know-spot as well and someone who can understand the situation from programmer's perspective hopefully. Then I can see a case being made. once again not a lawyer but an interesting question, would love reading other replies to your comment. also for what its worth, you can sue any company for X,Y or Z. The question worth asking is if you can win such lawsuit. Personally I believe it might be hard but not impossible but for all practical use cases it might as well be but the only answer can probably be found in court. I am just guessing at this point.