9 ms·
Incident with GitHub Actions, API requests, Codespaces, Git operations, Issues
- thomassharoon 5y agoPull review comments and approvals as well
- jetpackjoe 5y agoThe github.com homepage, as well as api (via `gh`) are not working for me either.
- niel 5y ago> The github.com homepage Only while logged in, it seems.
- jetpackjoe 5y agoTheir status page is reflecting the new outages. Good on GitHub for actually updating that quickly.
- Xarodon 5y agoPushing to repos is also not working
- sinkensabe 5y agosame here
- kitten_mittens_ 5y agoCan't push changes at the moment.
- deleted 5y ago[deleted]
- timeimp 5y agoIt’s not DNS There’s no way it’s DNS It was DNS
- rvz 5y agoHere we go again. GitHub going completely down at least once a month as I said. [0] So nothing has changed. That is excluding the smaller intermittent issues. Let's see if anyone implemented a self-hosted backup or failsafe just in case. Oh dear. [0] https://news.ycombinator.com/item?id=30149071 https://news.ycombinator.com/item?id=30149071
- bastardoperator 5y agoThe entire point of git is that it's decentralized, lol. If I've cloned locally like millions of people do daily, I have a backup.
- rvz 5y ago> The entire point of git is that it's decentralized, lol. No-one here is criticizing git itself. That is not the point. It is GitHub that is defeating the whole point of it all, once their hosted central server goes down. The majority of these projects went all in on GitHub, including using GitHub actions, npm packages, hosting their whole website, etc hence as soon as it goes down, they can't push or update anything; especially if it was very urgent. It has become a giant single point of failure for nearly everything. There is a reason why the Linux Kernel, Mozilla, Qt, Chromium, GNOME, ReactOS, etc self-host their own repositories and have fail-safes repositories if Github goes down and becomes unreliable.
- uplebian 5y ago> It is GitHub that is defeating the whole point of it all, once their hosted central server goes down. server != service assuming its a distributed service vs one server for a multi-billion$ company also group of humans built this service, so its not gonna be perfect :shrug: companies that use such tools and in trust all the business process to a provided service and do consider an event like this is a blocker should build in contingency plans or accept that there is no real 5-nines of availability more like 90-98%
- 5y ago
- Wavelets 5y agoWhew, glad I decided to scroll HN right now. I've been puzzling over why I'm getting "! [remote rejected] master -> master (Internal Server Error)" as well while trying to push and decided to take a break.
- m3nu 5y agodito
- forgingahead 5y agoIt's been like that for at least 6 hours, randomly appearing. I would take a pause and try again and then it would work, but now it's definitely much more persistent. Guess it's time to go play some video games.... https://xkcd.com/303/ https://xkcd.com/303/
- cik 5y agoAlso yesterday depending on where you were in the world.
- 5e92cb50239222b 5y agoHere you go: $ while ! git push my; do sleep 1; done Works for me eventually, although commits do not appear in web interface (they do in the actual repository).
- forgingahead 5y agoThanks but no thanks - no way am I doing anything to my core app repos when the repo host is fritzing out. This is one of those moments to go for a walk (or bed, depending on your timezone).
- TimWolla 5y ago-f does not sound like a good idea to me in a script like that.
- 5y ago
- candiddevmike 5y agoThis is causing actions jobs to hang after completing, consuming precious minutes. I don't think I've ever seen a refund when this happens, so I recommend everyone check their jobs and cancel them for now.
- WFHRenaissance 5y agoLooks like the drinking started early at GitHub... good on them!
- PeterBarrett 5y agoOne of our systems runs AWS code repository in parallel to Github and builds are triggered from there (but not in us-east-1). Time to migrate the rest of our systems to having that fallback.
- lebski88 5y agoIt's almost the same time as their incident yesterday too. Although today the scope is wider - yesterday it was Webhooks and Actions. Today core git is broken as well as the APIs.
- pm90 5y agoYep. I hope they post an aws style postmortem… this is kinda ridiculous (although I do empathize as an ops person). Webhooks breaking broke all of our pr bots bringing development to a standstill yesterday; today everything seems f’d.
- alexambarch 5y agoI'm unable to even sign out. It gives me a 500 and then drops me right back at the homepage on a refresh.
- arpinum 5y agoThese incidents have to hurt Azure's brand value. It's a monster task to run something as big as GitHub, if they ever get it stable it will lend a lot of credibility to Microsoft's cloud skills.
- jaywalk 5y agoI don't consider this a reflection on Azure at all. It's really just a reflection on GitHub under Microsoft's leadership.
- jamil7 5y agoEh, I'm no Microsoft fan but it used to have issues before the acquisition too. I can't really remember if it was better or worse.
- ryanbrunner 5y agoThere's not really all that much pointing to an infrastructure level failure - it's possible, but it's just as likely it's an application-level failure somewhere in Github's code. The API is returning 500s and not 503s and the failure is relatively quick, so it's not obviously a server outage.
- kortex 5y agoIt's yellow lights across the board, literally nothing is green. That's usually indicative of some sort of software infrastructure level failure or cascade failure, not an application-level failure, which usually manifests as one or two specific services going down (depending on how you define "infrastructure" and "application" - with IAC, arguably the software defined infrastructure _is_ an application). I doubt its a physical hardware issue. It's rarely hardware (except when your DS catches on fire). No red lights, so it's probably not something catastrophic like that facebook DNS SNAFU, but it definitely smells infrastructure- or deployment-scoped. Like either small DNS issue, or some load balancers are sending traffic to servers which cannot handle it programmatically (schema change?) so they are barfing.
- RapperWhoMadeIt 5y agoDo they regularly publish post-mortems after their repeated incidents? Might be interesting...
- samgranieri 5y agoI think they usually do, especially for the hairy issues.
- svnpenn 5y agoI cant even comment on issues...
- bombcar 5y agoIt's intermittent, I was able to get a push through eventually, and am now hung trying to convert a draft PR to ready for review. It took many tries to get to draft. I'm probably not helping by repeatedly trying, but I don't want to forget this PR. Yay it finally went through.
- Saig6 5y agoI'm able to occasionally push commits, but PRs aren't picking up the update or rerunning CI
- jhugo 5y agoIn Asia I've been having problems for almost 12 hours now (both locally and from our CI/CD which is in a different country). Also had similar problems on Tuesday.
- deleted 5y ago[deleted]
- bloopernova 5y agoAnd of course my developer teammates are still trying to merge PRs. I don't care that it works "some of the time"! Don't mess with the repos when the repo host is having seemingly random issues.
- fritzo 5y agoFor example: while actions are down, branches can be merged without ci tests passing, even for protected branches. This just happened on one of my repos.
- intsunny 5y agoWhew, outage timestamps in UTC. Now I won't have to know what time is it California, and if California currently has PST, PDT, PTSD, etc
- everfrustrated 5y agoDoes anybody else remember when GitHub's outage page used to have little graphs showing downtime? Eventually they took it down as their outages were just too often. GitHub has _always_ had terrible uptime. It's a great product - wish something would change but it seems cultural at this point.
- pythux 5y agoI have no idea if this is remotely close to reality but, what if, their culture of breaking things and bad uptime is what allowed them to move fast and build a great product in the first place?
- hn_throwaway_99 5y agoGitHub was founded in 2007. They were acquired by MS years ago. They should be well beyond any startup culture of "move fast at the expense of reliability".
- pythux 5y agoI don't disagree with this, they could/should have transitioned already. But for one, cultures are hard/slow to change. And second, as an example, Facebook had the motto "move fast and break things" until 2014, and by that time they also were beyond the startup phase(), so this kind of culture is not only for early days. () They were founded in 2004, that's 10 years in. By that time in 2014 they had 800M+ monthly active users and $12 Billion revenue; and they had this culture internally until this point.
- hn_throwaway_99 5y agoFacebook is a social media app that hardly anyone (except for advertisers) pays for. GitHub is an enterprise product crucial to tons of businesses. Cultural comparisons between the two really shouldn't apply.
- 5y ago
- avar 5y agoI'm finding that pushes do go through eventually, this is probably grossly irresponsible, so I don't recommend its use, but I remembered I had this old alias to "push harder" in my ~/.gitconfig: [alias] thrust = "!f() { until git push $@; do sleep 0.5; done; }; f" I've done a few pushes so far, and found that it's going through in <10 tries or so.
- hackandtrip 5y agoAdd some kind of exponential backoff to be a good citizen!
- 5e92cb50239222b 5y agoIt's fine. Maybe it will force them to finally start paying attention to the quality of their work. If crap I'm writing for a living was misbehaving that frequently, I'd be sweeping the streets by now (or doing some other work that's actually useful to society).
- ctvo 5y agoIt's OK to be frustrated since we rely on GitHub so much, but this is unkind. Software is complex. GitHub operates at a scale few of us work at. There are people at the other end doing their best traversing complex internal systems (organization and tech). I would argue GitHub has done more for societal good than most tech ventures, by the way.
- BukhariH 5y agoPeople tend not to be very kind when any product they pay for goes down. At the end of the day - our companies also have people that rely on our software working in order to do a lot of societal good.
- ilkkal 5y agoSure, but it’s incredibly naive to see gh having problems and go “they must not know what they are doing”
- deleted 5y ago[deleted]
- fishywang 5y agoThey just had a (smaller) outage yesterday. At first I thought it's yesterday's incident finally got enough points on hn.
- mr90210 5y ago
- mml 5y agozenhub appears to be having issues as well (can't load ticket at all) due to their GitHub integrations I assume.
- cedric 5y agoI downloaded a GitHub repo from Software Heritage [0]. I searched and found the repo was in the archive. Software Heritage saved my day. [0]: http://archive.softwareheritage.org/ http://archive.softwareheritage.org/
- Sydneyco 5y agoWhy is GitHub having so many issues recently? do you think it's due to the recent events?
- can16358p 5y agoAt some point GitHub main page 500'ed for me. The problem is probably somewhere down to the core, not at something isolated.
- jakub_g 5y agoAt least one good thing about GH is that while things break, the status page is updated relatively fast compared to other companies, when all HN knows about outage for 1h+ until it's acknowledged.
- deckard1 5y agoTwo days they have been down now. Github has, by far, the worst uptime of any critical service I've seen going on multiple years now.
- soraminazuki 5y agoAh, so this is the reason for the mystery failure I encountered with GitHub Actions. My job just failed without emitting a single error message.
- anarsdk 5y agoya’ll do know Git is a distributed VCS right? it’s ok for the the remote to be offline.
- lambda_dn 5y agoThis is why you should have your code on multiple remotes, i.e. Azure DevOps, Git labs, self hosted git server.
- i_like_waiting 5y agoWow, suddenly staying on-prem with old rusty Jenkins is not so bad. (It has its issues, but at least I had better service levels in last 12 months)
- orf 5y agoYou have to use Jenkins though.
- anunay_i 5y agodo they publish postmortem's? gist.github.com was down too for sometime