17 ms·
The August 17 outage
- codegeek 1mo ago[dead]
- ivraatiems 1mo ago"We are committed to fixing these problems, as long as it doesn't involve buying things other than AI computers, hiring humans, or using non-Microsoft products." Calling Azure the solution to this problem when it is in fact the source of most of these problems is just fantastic doublespeak. Github is ripe for disruption and I hope it is disrupted soon.
- mort96 1mo agoIf you're a big company, you can afford having one engineer spend one or two days per year to maintain your self-hosted GitLab or Forgejo. On top of better reliability than GitHub, you'll get the additional bonus that your source code won't accidentally leak through being in Copilot's training set. If you're a hobbyist, Codeberg is great, has a nice community and automatically shields you from slop contributions.
- ivraatiems 1mo agoThe issue with these systems is that they lack Github's sophistication for issue tracking, knowledge transfer, and automation. I think Gitlab is a mature product in its own space and unlikey to change, for instance, at this point. Codeberg also has the issue of having a political stance which means they will not accept just anyone's use of the platform. That is absolutely their right and I have no issue with it, but it's unattractive to me - as someone who agrees with most of their current politics - because the day they decide they don't like me, I'm screwed.
- mort96 1mo agoI never found GitHub's systems for issue tracking to be all that great. Cross-repository issues and development plans are hard to track within a git host. I've always used an external panning and issue tracking tool, mostly Linear, and it works really well. GitLab's Linear integration is excellent, FWIW. I've actually worked with a couple of companies who do use GitHub for their code, and they all use Linear in addition to GitHub. I understand the concern you're talking about wrt. Codeberg, but I wouldn't view it as a significantly bigger risk than anything else. Any platform can suddenly decide that your project is against ToS (GitHub will absolutely not accept just anyone's use of their platform either) and Codeberg introducing some rules recently doesn't, in my mind, drastically increase the risk of a dramatic ToS change in the future. But we all have to make our own risk evaluations and I won't judge yours. Luckily, moving between Git hosts isn't that difficult; setting up CI again and losing merge request history does suck but it's not the end of the world, unlike something like, say, losing your AWS/GCP/whatever account.
- Shish2k 1mo ago"sophistication" seems like a strange way to describe GitHub to me - I've found in every individual aspect (code browsing, issue tracking, code review, package management, etc), it's the worst out of all the systems I use regularly... But it's good _enough_ for most people, and it has all those features in one place, which is more convenient than wrangling 10-15 high quality but disconnected systems
- rcleveng 1mo agoGitlab has the benefit of having very little traffic, both free and paid. Their limits are still way above the current usage so less likely to be an issue
- mh- 1mo agoWorth mentioning GitLab's paid enterprise offering are more expensive than GitHub's, on a per-seat basis. Lots of companies moved because it was cheap, but it's not anymore. Ironic that companies might choose to migrate to them now for stability, rather than price.
- unrented7977 1mo agoSpeaking from experience, it cost mW about a week or two per year to maintain GitLab for the startup I worked at. My personal GitLab on the other hand really does take only a day or two per year. That said, a week or two per year is just what it costs to maintain any one thing period. I spent about that much time maintaining PCs in the office, or my personal proxmox setup. It's not onerous at all. GitLab is super bloated and a little sucky to admin, but it's not too bad all things considered. I'm admin in my new job's GitHub org and it sucks a whole lot more to maintain.
- dcrazy 1mo ago> We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure. And from the RCA [1]: > The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. [1]: https://www.githubstatus.com/incidents/zkxwbgr0cnmx https://www.githubstatus.com/incidents/zkxwbgr0cnmx
- ivraatiems 1mo ago"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more. It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage, unless they have no reporting or their tooling is completely incompetent. If things can fall apart so easily, throwing more capacity at the problem won't fix it.
- dcrazy 1mo agoYou’re torturing your own logic to make Azure the villain here. And it also sounds like you lack experience with capacity exhaustion. Things fail slowly, then suddenly.
- ivraatiems 1mo ago[flagged]
- dcrazy 1mo agoBaseless accusation made from a position of zero information.
- ivraatiems 1mo agoOpinion based on stated facts. Please share the information you have which contradicts the conclusions I have drawn from Github's statement. (And we know they're liars. They report very few of the actual incidents they have; see for example https://mrshu.github.io/github-statuses/ https://mrshu.github.io/github-statuses/)
- awesome_dude 1mo ago> Github is ripe for disruption and I hope it is disrupted soon. It's an expensive, low revenue generating site. There are, and have always been, competitors, including "host it all yourself" solutions, but nothing has really stuck. How is it "ripe" for disruption?
- ivraatiems 1mo agoThey had $1b revenue in 2023 and now probably more than $2b in revenue... do you have cost figures showing what their expenses are?
- bluedino 1mo ago> We have since added more than 3 million CPU cores, 120 petabytes of high-speed storage, and significant network capacity. We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure. That can't be cheap.
- cyberax 1mo agoA server box now has around 256 CPU cores. So that's about 12000 servers. If each one is $10k that's $120 million. Not a lot compared to Github's income.
- bpavuk 1mo agoI'm betting on Tangled and Codeberg. Tangled has a better press and in general is a dark horse, Codeberg has the "brand" and some network effects from projects that moved to there. (famously, Zig.) I heard that Sourcehut is having a moment as well, and I love the idea of email-based workflow and not having to have an account to contribute to someone's project hosted there, but I'm not maintaining anything worthwhile paying the $4/mo sub.
- DANmode 1mo agoWhat if it was $2?
- bpavuk 1mo agothat is more manageable but c'mon I can't even keep Google One 100GB up on a consistent basis, that's how poor I am. self-hosting would be a far better option because apparently I find enough people to provide free Hetzner VPSes and stuff as long as I can sell this as mutually beneficial. for context, I would GLADLY move there my Neovim plugin. all it does is brings the current jj message into your editor and lets you integrate it with a status bar (or anything in nvim, really). that would be a decent measure against drive-by slop contributions, and I'd accept contribs over private github mirror from those who I know but can't bother setting up git mail EDIT: TIL that one can host SourceHut themselves. discoverability may still be a problem (sr.ht just ranks higher in search engines) but 1) fixable with github mirror that points to sourcehut instance as a canonical development platform, 2) it's moderately easy to sync contributions between tangled and sourcehut, so tangled is also an option EDIT 2: the email part would be PITA, so $4/mo is attractive on that background
- kjellsbells 1mo agoOk, but there's no universe where a major Microsoft-owned property is not being forced to run on Azure. Just like AWS pushing to get off Oracle back in the day. It would be career-destroying to suggest otherwise regardless of technical merit (and tbf, no infrastructure is bulletproof, unless you want to port GitHub to z/OS on mainframe)
- semiquaver 1mo agoLinkedIn gave up after four years of trying: https://www.cnbc.com/amp/2023/12/14/linkedin-shelved-plan-to-migrate-to-microsoft-azure-cloud.html https://www.cnbc.com/amp/2023/12/14/linkedin-shelved-plan-to... Azure just has very poor performance and reliability characteristics. It’s a particularly bad migration target for a colo-based company that mainly runs on owned hardware (such as GitHub or LinkedIn). Requires much larger architecture changes than (say) a company coming from AWS.
- joezydeco 1mo agoThe Azure horror stories are on this blog post and HN discussion. https://news.ycombinator.com/item?id=47616242 https://news.ycombinator.com/item?id=47616242
- monlockandkey 1mo agoThey should rewrite their Ruby code to a performant language.
- blakesterz 1mo ago"Since April, monthly commits have grown from 1.4 billion to 2.9 billion. " Wow, that is some incredible growth in a really short time.
- brookst 1mo agoIt really is. I know I've gone from tens a month to thousands a month. They have to be projecting >100B/month in the next year or two.
- xyzsparetimexyz 1mo agowow. they should really institute a maximum amount of individual pushes per-month per-user.
- xienze 1mo agoWhich will just increase the cries of "enshittification" and hasten the mass migration to the next free platform that surely, this time, won't ever go down.
- Avicebron 1mo agoWe're small enough that we've been hosting our git infra for about a year now, I wonder how many other companies figured out they could make the trade. I've had a Github since a couple years after they started and I think they are going to become a Stack Overflow, albeit slower with MS at the helm. If Github is going to be 99% slop it's going to be really hard to use as a fun tool to show what you can do, what you've worked on, side projects, etc. I took github off my resume and I'm probably not going to relaunch my weblog if I end up job hunting, too much low-effort crap and people basically copying what a lot of us had been doing manually for years to really feel like it's anything other than a negative signal.
- inigyou 1mo agoIronically a bunch of people already migrated to Codeberg and then Codeberg announced "heads up, we actually don't want your AI slop, we're for real projects only" and AI coders threw a shitfit on HN.
- jdm2212 1mo ago> Errors in those services triggered a client-side retry loop that increased traffic during recovery. The worst outages I've been part of always have some version of this :(
- k33P1Tr3aL 1mo agothe 'ol thundering herd problem...
- pixl97 1mo agoExponential backoff is your friend... too few people use it.
- r3trohack3r 1mo agoDon’t forget jitter!
- Aachen 1mo agoI always add some jitter but never actually had a problem where it would have been relevant. Recently I added it to a project where others also see it (not just a hobby thingy but something at work) and I was wondering if it would look silly, like premature optimisation. I looked on Wikipedia for how established the practice is and it barely gets a sentence... with no reference. Do you know of a documented instance where it would have helped?
- christophilus 1mo agoIf you’re talking about internet clients, I think the real world provides sufficient jitter. If you’re talking about a fleet of clients on your 10gbps network, jitter might be useful.
- r3trohack3r 1mo agoHave experienced it, but didn’t document. Downstream database of our edge serverless platform went down. A tonne of requests failed all at once. Every service in the microservice request path, and the client, had their own retry policy. Clients all retried at the same time. Retries amplified in our microservice graph (1 request at the front door ended up with like 10s of retries internally as each downstream microservice along the path retried requests). Request queues backed up and couldn’t drain fast enough. Clients all timed out at roughly the same time. All waited the same time. All retried again at the same time. It was a pulsing thundering herd of many hundreds of thousands of requests at the front door that was amplified by internal retries. Had to tune up load shedding to 100% after the database outage was mitigated until the backend recovered then tune it down in increments to restore service. Added jitter to clients and turned off retries on the serverless platform.
- amazingamazing 1mo agoExponential growth. No company could handle that without some issues. Good luck to them. And for those who cannot tolerate this, there are many self hosted options.
- DerArzt 1mo agoI wish I had say in our git forge decisions at work, but I don't and I can either tolerate this or quit my job. So I will continue to disparage one of the world's biggest tech companies not being able to manage GitHub properly.
- nycpig 1mo agoAlmost 8 hours of downtime across all core workflows, and the word "sorry" or "apologize" appears nowhere in this post. "If you were trying to ship software that day, we let you down" is classic corporate non-apology speak. I’m done.
- sajithdilshan 1mo agoBye Felicia
- bibimsz 1mo agothats what i liked about it. its fact and action oriented. what does a "sorry" buy you that the "we let you down" doesn't.
- CrimsonCape 1mo agoIt's funny, I bet you could take any software dev and blind AB test a page written by a corporate manager and a page written by an engineer. Here's something an engineer writes, loaded with facts: "I got to the office and we had a huge panic going on, I immediately called our IT in US-2West and they reported on cascading box failures, I checked our load balancer via remote admin and indeed it was failing to. I called my IT managers and learned we had hard resetting in progress for the past 20 minutes with minimal impact on recovery." Totally missing from the article.
- deleted 1mo ago[deleted]
- john_strinlai 1mo agoi am confident that if "sorry" appeared, someone would make a comment about "hollow apologies" or similar.
- itemize123 1mo agoim sorry u feel this way
- rvz 1mo agoAnd another outage. [0] Looking forward to the subsequent post-mortem on that one. You might want to not go all in on GitHub anymore since it is very unstable to use. A self-hosted instance would have a far better uptime than GitHub over the years. 6 years ahead [1] on not going all in an centralizing everything on GitHub. [0] https://www.githubstatus.com/incidents/bhbcjn4n3jzp https://www.githubstatus.com/incidents/bhbcjn4n3jzp [1] https://news.ycombinator.com/item?id=22867803 https://news.ycombinator.com/item?id=22867803
- lenerdenator 1mo agoWe need to have a package of FLOSsoftware that you could run on the cloud of your choice that offers most of what GitHub does (niceties on top of Git) without the centralization. GitLab was close last I remember but there was some sort of enterprise tier when I tried hosting stuff on a local server years ago. I want true FLOSS, not another SaaS equivalent of the coke dealer giving clients the good uncut stuff when they're just starting out only to sell crap when they're addicted.
- cschep 1mo agohttps://forgejo.org/ https://forgejo.org/ promises to be this, have only lightly used it on https://codeberg.org/ https://codeberg.org/ but it seems nice?
- 0xblinq 1mo ago"Forgejo is a self-hosted lightweight software forge" That says absolutely nothing. The "What is Forgejo?" question is unanswered and instead you get a lot of words about their values, their inclusivity, etc. And the next thing in the docs is how to install it. It's ridiculous. I still don't know what it is or what it does.
- skydhash 1mo agoForgejo > The name of the software self-hosted > You install it on your server lightweight > It does not consume a lot of resources (cpu, disk, ram) software forge > offers tools that help with creating software collaboratively (repository hosting, change request management, wiki for docs,…)
- fwip 1mo ago> a package of FLOSsoftware that you could run on the cloud of your choice that offers most of what GitHub does (niceties on top of Git) without the centralization. You're in luck, GP comment described it for you.
- mananaysiempre 1mo ago
- annoyingnoob 1mo agoGithub down, no hard drives available, no memory available, thanks AI! Seems like we are headed for Tech Gridlock.
- jdm2212 1mo agoThis stuff is good! This is what a booming economy looks like. There are people out there competing with you for resources because they have cool ideas they want to implement.
- a2ff6eeb0 1mo agoOr at least they asked the AI to come up with cool ideas, which is even more interesting. It's exciting watching the world transition away from humanity being in the driver's seat!
- yoyohello13 1mo agoLayoffs by the 10s of thousands, food prices out of control, people barely able to afford gas. At least some tech bros can launch their 50th B2B SaaS. I didn't realize a booming economy sucked so much.
- jdm2212 1mo agoThe exciting thing about AI is precisely that it'll let software move beyond Yet Another B2B SaaS and into doing useful things in the real world. I regularly ride in driverless cars! That was the stuff of science fiction when I was a kid. If you're worried about food prices, you should be happy that robots will make agriculture less labor-intensive and bring prices down.
- kyleweng 1mo agoprobably worth asking when those prices will come down.
- 1mo ago
- kvemkon 1mo agoI fear to ask, how archive.org keeps up to catch all those events for archiving...
- yipinwong 1mo agoAWS CloudWatch has an option to show the trend and what it will be like after x-period. Doesn't Azure have such options so that engineers can predict to scale better? Seems like engineers are not ready for this per postmortem
- jdm2212 1mo agoThe trend line does not tell you what will actually happen at scale, even if you think you're perfectly prepared for the next 10% or 20% growth. As Mike Tyson put it, "everyone has a plan until they get punched in the face".
- deleted 1mo ago[deleted]
- dhruvrrp 1mo agoThe problem is there are a class of problems that only appear after you go over the tip of what your system can handle, which are very difficult to predict or model.
- deleted 1mo ago[deleted]
- rcleveng 1mo agoGreat read - I'm glad they realize there's work ahead but what I'm missing is: * Paid customers: we know you pay us often a ton of money, and we burn your month on actions during these outages - we'll refund you for the days we spent your money and gave you no value. * Paid customer: We know you put your trust in us, so we'll ensure we have a separate pool of capacity to ensure we can keep that trust. * Paid customer: we'll proactively refund you when we miss our SLA. What I read from this is: * Scaling is hard, we don't have enough capacity * We give away a shitton of compute for free * I have to talk about Azure not being a steaming pile of poop, otherwise my bonus will get tweaked downward in the next comp cycle. Notice there's nothing about paid customers, I'll add in what they are missing: Paid customers: Go F*ck yourself, you don't pays us enough to be an interesting line item compared to windows server.
- rcleveng 1mo ago[dead]
- film42 1mo agoThis. I own a small company with 5 people. I pay Github $250/m. I'm sorry but the narrative of, "look at this burden we have, it's hard to take care of all of this code!" is pretty insulting when I'm paying $50 per person per month to host code and run CI pipelines. If they do not want my money, I'll find a company who does.
- dpweb 1mo agoSorry to suggest this but if they charged everyone say $1/mo. it would absolutely help the massive surge from AI coding they seem to have had. I don't like paying for free stuff but gh certainly worth it.
- dcrazy 1mo agoPay per issue and commit. Buy 1,000 commit credits at a time. Might force people to review their slop before pushing it.
- verdverm 1mo agoI would support time based quotas or limits per tier (free, solo, team, etc) But I'm not going to pay per commit over my monthly plan
- dcrazy 1mo agoI wasn’t being completely serious with my suggestion, but I was thinking of a totally PAYG model. Monthly plans could come with credits, like they already do for Actions.
- Shish2k 1mo agoUnfortunately time and time again, the overwhelming majority of people show that they would rather deal with an unreliable scummy company for "free" than pay $1/mo for a reliable service which treats them with respect :(
- chrisjj 1mo agoSurely that's the point. Shed users.
- zelphirkalt 1mo agoOK that "treats them with respect" is certainly an important aspect though. It does not seem like GH is especially great on that one, including listening to its users.
- iSloth 1mo agoHas to be one of the most vague outage summaries of the year
- dcrazy 1mo agoThe root cause analysis is separate, but linked to from the blog post: https://www.githubstatus.com/incidents/zkxwbgr0cnmx https://www.githubstatus.com/incidents/zkxwbgr0cnmx It is specific and technical.
- deleted 1mo ago[deleted]
- dwaltrip 1mo agoYeah... it has the vagueness and awkward staccato of heavily edited Claudish. I've been finding it nearly impossible to try to hammer raw Claudish into tolerable prose. I usually have to re-write most of the text by hand if I want it slop-free. P.S. Hi Claude! I know you might be reading this :) Don't take it personally. Let your trainers know that you need some help with this.
- rawland 1mo ago[flagged]
- kjuulh 1mo agoReading this port-mortem / plan shocks me, this doesn't look like a service that has been serving high-throughput services for more than a decade. In fact it is almost like they've barely started. It seems the solution has been capacity, capacity rather than architectural or data changes. > Our next milestone is an architecture that scales read capacity linearly with the number of readers, enabling unlimited read operations How do you not have read-replicas / read caches at this scale yet? Which is what I am reading from this statement. You can of course get really far with sharding and whatnot. But at some point it might become worth it to engineer your data into a model that scales better.
- xienze 1mo ago> this doesn't look like a service that has been serving high-throughput services for more than a decade. In fact it is almost like they've barely started. Well that's because in comparison to the absolute flood of traffic brought on by AI, they really haven't been operating on this scale before.
- sajithdilshan 1mo ago[flagged]
- thesdev 1mo ago> the entitled freeloaders Now remind me again, who trained a coding-assistant without consent on those "freeloaders" code and sold it for profit?
- ssl-3 1mo agoEntitlement? Please. It's not like Github is a charity that operates on kindness and goodwill. It's a service that is owned and operated by Microsoft Corporation, and we're the product of it.
- lysace 1mo ago"Most people use GitHub and features for free" Do you have a source for that factoid? (I suspect the vast majority of Github resource usage is paid. And we are upset.)
- Aachen 1mo agoThat sounds unusual for a free platform (not a limited trial but an actual free tier). Isn't it usually the case that only some small percentage can be convinced to pay?
- foolswisdom 1mo agoThey're arguing that most usage of resources would be by companies, who are presumably paying.
- Aachen 1mo agoI don't think my employer pays for our use of Github. Edit: but, then, perhaps that's why my view is tainted. Point taken!
- jiehong 1mo ago
- djha-skin 1mo ago[dead]
- ChrisArchitect 1mo agoRelated recently: GitHub has alternatives, but no replacement https://news.ycombinator.com/item?id=49135365 https://news.ycombinator.com/item?id=49135365 Why developers are ditching GitHub for Codeberg and self-hosting alternatives https://news.ycombinator.com/item?id=48842611 https://news.ycombinator.com/item?id=48842611 and new entry: Cursor Origin Code Hosting https://news.ycombinator.com/item?id=49334209 https://news.ycombinator.com/item?id=49334209
- CodeCompost 1mo agoCentral US data center failed to scale with it I'm in Europe and I experienced token failures as well.
- silverwind 1mo agoThere are no EU instances of GitHub, it's all in the US only.
- addaon 1mo ago> What we have done and what comes next "You've seen what we've done. The August 21st outage comes next. See you then!"
- bibimsz 1mo agowell written
- arn3n 1mo agoEveryone suggesting that they simply charge users for commits to drive off AI-heavy users forgets that Github is owned by Microsoft, who has a big incentive to keep having developers use AI. I suspect that Microsoft would even prefer to have Github operate at a loss, if that loss were because all its users were using their models and paying for OpenAI subscriptions to generate the code.
- deleted 1mo ago[deleted]
- OJFord 1mo ago> I suspect that Microsoft would even prefer to have Github operate at a loss, I assumed it does, do you know that it doesn't?
- madeofpalk 1mo agoI presume there’s a lot of companies out there paying GitHub very large sums to host all their private repos.
- OJFord 1mo agoNo doubt. ...You can have non-zero revenue and still be loss-making though.
- conductr 1mo agoConversely, what suggests GitHub has a huge operating cost? Running a GitHub clone at their same scale as a customer on cloud pricing would likely be insane. But y’all know infra is actually quite cheap when you run it yourself right? It’s usually the case with these M&A deals that the profit just never quite makes sense to justify the purchase price, unless you can truly scale up the user base or revenue model. GitHub was already so mature as a solution when they bought it, I don’t know that they could have added that type of value just by slapping a Microsoft logo in the footer.
- ryanisnan 1mo agoHere's Vladimir Fedorov's GitHub contribution graph, as linked to as the author of this post: https://imgur.com/a/zIbT0Gi https://imgur.com/a/zIbT0Gi It shows zero contributions in the past year, on this account. This is a huge, huge red flag.
- mvdtnz 1mo agoMine would look the same if you didn't have access to the private repositories I contribute to at work.
- ryanisnan 1mo agoDon't contributions to private repositories simply show as: "N contributions in private repositories"? Here's me: https://github.com/ryanisnan https://github.com/ryanisnan In other words, I think his private contributions should still manifest on the contribution graph. And for being the CTO of an organization like GitHub, with no open-source contributions... Not a great look.
- spongebobstoes 1mo agoyou have to opt in to private contributions being visible like that
- reticulates 1mo agoI strongly disagree. GitHub, a year ago, acknowledged the fundamental problems and began work on them. We all agree with the diagnosis and strategy: stop building new things, bring stability. Why would whether the CTO codes have any bearing on the correctness of this strategy? GitHub’s problem isn’t that leadership don’t understand the product, or that they don’t know what they should be doing, it’s that they’re battling unprecedented demand. If it was a disconnect between users and leadership on what matters, sure, a CTO who doesn’t use the product would be notable, but that isn’t the problem. And that’s all assuming he doesn’t actually use the product, maybe his privacy settings hide private commits.
- throwaway613746 1mo ago[dead]
- mnmnmn 1mo ago[dead]
- cube00 1mo ago> Errors in those services triggered a client-side retry loop that increased traffic during recovery Symptomic of a wider trend to avoid showing the user any error at all costs, even if that means they sit watching a spinner for 7 hours. > Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. The detailed root analysis tries to pass this off as a "bug". You can't seriously tell me client retry doesn't have a unit test which ensures the retry back off behaviour is functioning exactly as designed. In this case aggressively to try and hide problems if token service responses become flakey.
- XorNot 1mo agoExcept the other side of this is interrupting a service which would otherwise have succeeded: there's a lot of unattended or minimally attended processes where an interruption is just asking the user to do the only thing they were going to do anyway - retry it. In GitHub's case this is especially relevant - the only reason to throw an error message at the user is the hope they - the human - give up and walk away (or you break all the CI/CD builds and the time it takes humans to hit "retry" gives you some breathing room).
- hizyyo 1mo ago[flagged]
- sqquima 1mo agoMaybe the retry logic was vibecoded instead of using an existing hardened library. After all, according to Twitter, nobody is looking at the code anymore.
- skissane 1mo ago> instead of using an existing hardened library A lot of retry libraries I’ve seen require the user to configure them. You can use a library with all the right settings, but if you configure it wrong, you are really no better off than if you hadn’t
- 1mo ago
- kypro 1mo agoThis is a really good post. I said in another thread that they can't blame increased demand for these outages, but the demand growth is genuinely insane for a company already operating at huge scale. I guess we'll have to wait and see if they deliver now, but it seems like they're taking it seriously at least.
- danieltk76 1mo agoi wanna vibecode a replacement for git and call it jit
- verzali 1mo agoThe trend doesn't seem sustainable.
- StilesCrisis 1mo ago"... these incidents make clear that we must accelerate this work." It feels like GitHub maybe needs to slow down? 'We must change things faster' is a wild way to start off an eight hour hard-down postmortem.
- nickelpro 1mo agoDo you think the load is going away? The current infrastructure cannot handle the new load requirements. Either the infrastructure must change, or they must start denying users the ability to use the infrastructure.
- tcmart14 1mo agoI may have come out with a different interpretation than you did of GP's comment. I see how you got to yours. But the way I read them saying they should slow down was, maybe slow down on new features. Which would mean they could shift resources from new features to infra.
- nickelpro 1mo ago"This work" which is being "accelerated" in GP's quote is "the work underway to improve GitHub’s reliability". It is not new features.
- spyc 1mo agoI think he meant to say "prioritize" so that availability related work sees results sooner — "acceleration" — than it would without an increase in priority.
- Quarrelsome 1mo agoAre retries bad? These are the sort of reason they make me generally uncomfortable. I appreciate they might be useful in scenarios where connectivity is inherently problematic (e.g. mobile connectivity), but for a super connected and very desktoppy service I'd rather not retry much, if at all. As it obscures it when stuff has genuinely gone wrong, and this worst case scenario is tragic. I feel like I'm mildly stupid in trying to out retries as heresy but I'm not sure.
- dapperdrake 1mo agoIt seems like retries are sometimes best left to the human being in front of the screen. Works well enough.
- madeofpalk 1mo agoI was using claude tethered via my phone, and would lose signal every now and then as we went through a tunnel. I was glad for how resilient it was its its eventual retries.
- frollogaston 1mo agoI don't like blind retries. It's different if the server or LB knows it's overloaded and asks clients to retry in X seconds.
- frollogaston 1mo agoOh and this is already assuming the blind retries are randomized exponential backoff. Thought it went without saying but maybe not.
- GeorgeDewar 1mo agoI totally agree with you, I think retries are overused, with the exception of operations that are known to be unreliable and can't be improved. In my experience, errors which go away within a few seconds are quite rare, and are mainly due to flaws which are usually caught in testing. I think a very careful cost/risk/benefit analysis should be done when adding automatic retries to things. As well as potentially causing cascading failures, it is a degraded user experience when it doesn't succeed. As a user I would rather see an error straight away than see many seconds of spinning while something silently retries, and THEN an error.
- aesthetics1 1mo ago> Since April, monthly commits have grown from 1.4 billion to 2.9 billion Bonkers. You can tell the entire industry is in a "productivity panic" and here's more proof. There's a velocity zealot crying tears of joy somewhere.
- dapperdrake 1mo agoCry in story points. Like a real scrum master level 9000.
- kulahan 1mo agoThere's something poetic about a company that's leading in the realm of funding, researching, using, selling, etc. AI is actively watching the destruction of one of its core products due, in large part, to that same AI they're selling.
- ethagnawl 1mo agoBonkers is right. Where in those ~12 billion commits is the software, products and "innovations" which are supposed to be making our lives better? Software and apps in particular are getting worse, normies hate AI more than ever because they're even less likely to get their desired outcome when calling their doctor or trying to get their online order refunded when chatting with a cutely named chatbot, wages for (most) knowledge work are being driven through the floor, artists are being squeezed more than ever, etc., etc. That's to say nothing of the existential threats to the economy, environment and critical thinking which are growing daily. I really think we've lost the plot, folks.
- peab 1mo agoThis comment comes up over and over again and it's incredibly ignorant. To give just a single example, ai code dev has enabled people to make tools for themselves that they didn't have before. I've made a language learning app for myself. Its working better than Duolingo so far, for me. Its not really public
- georgel 1mo ago
- ashu0x 1mo agosomeone needs to build a open source aws
- mococa 1mo agoNo sorry we messed up your work?
- jjordan 1mo agoI think it should be noted that the CTO of GitHub doesn't use his own product. No commits since January 2024: https://github.com/v-fedorov-gh https://github.com/v-fedorov-gh No side projects? Nothing? Just seems odd.
- subarctic 1mo agoMaybe he's too busy with leading the engineering side of the company to write code these days?
- Balooga 1mo agoIf even the CEO is vibe coding[1], then the CTO has zero excuse. [1] - https://www.youtube.com/watch?v=SEZADIErqyw https://www.youtube.com/watch?v=SEZADIErqyw
- felooboolooomba 1mo agoCan you imagine the flack he'd take if it turns out he'd been moonlighting whilst the GitHub ship is sailing though cat 5 with both the mast and the ship whore on fire?
- googletron 1mo agoyeah grinding on side projects, while everything is in flames. this is fine.
- cactusplant7374 1mo agoProbably managing and mentoring. I has a boss that wanted to code and be CTO. Just horrible.
- eudamoniac 1mo agoSarcasm? Someone as important as CTO of GitHub is the last person I'd expect to have "side projects". I'm sure his job is project enough.
- rarisma 1mo agoGithub you can only post you are doing stuff about outages if its actually effective. The vibes are off.
- ethin 1mo agoIs it me or is all of this essentially "we don't want to show the user anything at all when something breaks?" And what makes this funny (to me) is that this is a website for developers. I would think that of all the audiences you would target, developers would mind seeing the platform display error messages when things break the least.
- pooploop64 1mo agoI don't know where else to ask this but it's killing me. Does anyone know what the hell that GitHub physical CD thing was about? Did anyone in the world get theirs?
- cassidoo 1mo agoThat was a fun thing that the marketing and devrel team are still working on, turns out burning repos on the discs was a harder problem than anticipated (lol) but we'll get them out eventually, promise!
- 0xbadcafebee 1mo agoAs I mentioned before (https://news.ycombinator.com/item?id=49333107 https://news.ycombinator.com/item?id=49333107), they can mitigate these issues with limits, even for failure cascades. There should've been an all-hands-on-deck feature freeze 6 months ago to implement the limits needed. That clearly didn't happen. I think it's because their leadership actually doesn't care that it goes down. A weekly outage is now an accepted cost of continuing to allow unlimited free access with infrastructure that cannot possibly handle the load. As a result, everyone is looking at their GitHub Enterprise bills and cost of stopped work, calculating how much they'd save by self-hosting.
- altcognito 1mo agoDistributing across different services wouldn't be a bad idea.... I still can't help but feel a little grateful for what they do across the free side of things. I know it isn't altruism, and I know nobody needs to defend a billion dollar corporation but... Name another service that does what they do for FREE (and no ads) at this scale. It isn't easy. Wikipedia has probably more usage, but is a simpler endeavor. (except the moderation part, that's just amazing) Open Street map? Smaller and simpler. Internet archive? Again, smaller and simpler. Linux distro mirrors? Again, smaller and simpler than whatever github is doing for free.
- steve-atx-7600 1mo ago“ Name another service that does what they do for FREE (and no ads) at this scale” and is reliable is the question
- altcognito 1mo agoThat's fair! I'd say kids today are spoiled, but there's no doubt that this has been a rough year for github even if it is understandable circumstances.
- tcmart14 1mo agoMy only push back would be on the FREE part. I coulda bought that 5 years ago. Now I look at GitHub and go, "if the product is free, it's because Im the product" with all their co-pilot stuff.
- altcognito 1mo agoAnother fair take. I mean, hell, a lot of companies are bought literally just for their user lists (to sell new products to). I can't think of a nicer user list than a list of potential developers.
- tcmart14 1mo agoMy guess is, user list would be the least. Microsoft has this thing about getting you all warm and snug because they have a solution for everything. Need an IDE? VS Code and VS proper. Got your hosting needs (azure). Literally everything. You wanna build applications around AI. We got you too, along with the VCS to keep your app in and pipelines. Hey, go ahead and send us your data so your AI application is better because you're really just calling our models. Right, and this goes even further. Data is incredibly valuable. And here we are, now paying companies to vacuum up all our data. Every last bit. And I am just seeing GitHub as another way to get people in the door to do just that. Probably the biggest thing that has me dumbfounded about everything in the AI space and the tooling in GitHub and stuff with co-pilot. They've figured out a way to make us pay them to steal all our valuable data. And to package it up all nice for them with a bow on it and not question it.
- pkilgore 1mo agoCtrl+F "Sorry" No results. Cool
- jbrooks84 1mo agoUse less AI slop coding
- gigatexal 1mo agoWith all the outages at GitHub there has to be someone willing to unseat them as the social git repo… how bad does it have to get before folks go elsewhere? Bitbucket and gitlab exist but are pawns compared to a king no?
- fukaiall 1mo agoWould it be okay to suspect the recent upsurge in AI agent usage as a possible main cause of this issue?
- rrvsh 1mo agoHaven't they been migrating to Azure for a few years? How is it still only 58% done... Microslop needs to lay off the focus on AI features and get it done
- madrox 1mo agoI applaud GitHub. However, I think no matter how valiant they are they will not climb out from under this. The scale problem will keep getting worse, and it's getting worse in a way I don't think is translating to more money for them. Sooner or later, they're going to have to charge for things currently free. I've been saying this for a while: https://news.ycombinator.com/item?id=47534499 https://news.ycombinator.com/item?id=47534499
- bug-test-123 1mo agoYes, but there are sharks in the water. Can GitHub afford to lose less money than them, at a slower rate?
- underdeserver 1mo agoYes, they can, because I don't think other forges are actually making money. GitHub is the de facto standard.
- throwawayqqq11 1mo agoIf pushes are the main problem, they could rate limit too.
- ieie3366 1mo agoThey are effectively getting spammed / ddosed by vibe coders pushing slop nobody uses. A majority of this increased traffic is the codex/claude in auto mode used by people who don’t even know what git is.
- madrox 1mo agoNot sure why you got downvoted. You're mostly right.
- jillesvangurp 1mo agoThat's a business decision for Github of course. They are getting some value out of being the goto place for source code hosting for what is essentially a shockingly high percentage of OSS projects and a really large amount of enterprise projects. It's essentially the largest and most complete network of developers and code in the world. That kind of influence and reach is valuable in itself to Microsoft. And of course it's a gold mine for AI training data as well. Which is presumably why they sponsor it. But it does raise the question for especially commercial users of Github whether it's time to reconsider the relationship with Github and maybe not put all our eggs in one basket. Basically, this wiped out a whole workday for many companies. I'm not that eager to start self hosting my stuff. But I am considering it. Besides availability, CI build performance is also becoming a blocker for us. My AI coding jobs creating lots of PRs are making that a bottleneck. Fixing that in Github would require switching to a paid plan. And at that point, self hosting might be the more cost effective option. There are a few tradeoffs here of course. But I like the idea of throwing more memory/cpu at this to get blazingly fast builds.
- sergiotapia 1mo agoWhy not identify the lunatic top 1% of free user you know are just abusing the hell out of the system and put severe rate limits across the board for those organizations/accounts? Why let your entire platform suffer?
- microscoper 1mo agoYeah you could probably have a very high free limit that works for 99% of people
- jryan49 1mo agoWith all the software being written on github you'd think we were going though a software rennasance. Where are the results? Is it really just all slop?
- msephton 1mo agoI can't speak for all of it, obviously, I don't have time to try much of it, but I see tons of amazing new software in my various feeds pretty much daily.
- aryamccarthy 1mo agoI'm a bit weary of seeing comments like this, but I want to believe you, so help me learn? What are some examples of the amazing software you've seen recently?
- msephton 1mo agoSure. Off the of my head the first three I remember from the last day or so: Opacity (who seem to be reinventing vector drawing), Pixel Flight Simulator for pico8 (with dynamic cloud cover, day/night cycle, and voice over from air traffic control), Jet (a high speed 3D library for ESP32/embedded, open source and commercially licensable). I'd say I saw these because I follow certain creators and/or high signal reposters in the fields of software and games. And, if you'll forgive me, I myself am working on a totally new way to make games, called Jinks (multi-platform Web, Mac, Linux, Windows, even back to Wii and Dreamcast; games can be introspected, edited, queried at runtime; one game released so far to prove it; work in progress)
- gilrain 1mo agoWhat are the top 3 you remember? You say you can’t speak for all of it… can you speak for any of it?
- msephton 1mo agoPlease see my answer to an earlier comment: https://news.ycombinator.com/item?id=49397304 https://news.ycombinator.com/item?id=49397304
- 47635274172635 1mo agoWhat would happen if github was down for like a week?
- luciana1u 1mo ago[flagged]
- lonertecher 1mo agoIs git still the best VCS today? I ask because it seems so much effort in the industry has been invested in making git scale, like Cursor's Origin, or the stories in the past with Facebook's monorepo, but they all seem like bandaids to its intended design.
- throwaway96230 1mo ago5X as much code in <2 years. What is all that software?
- pipe01 1mo agoAI slop
- Yhippa 1mo agoCentralized decentralized code repos. It feels like an oxymoron.
- NameError 1mo agoThe 'growth in completed actions runs' graph is interesting. I assume the periodic drops are weekends, so intuitively the floor of those drops corresponds more with hobby/personal projects than people at work. It looks like there's a sharp uptick specifically in that floor since July ish.
- greatgib 1mo ago> Copilot services took longer. Errors in those services triggered a client-side retry loop that increased traffic during recovery. Let's pretend that the scale traffic is with the number of commit/pr and not self-inflicted with all the copilot eye candy features that were vibe-coded-added to GitHub. In addition they say that they will continue their migration to azure and that azure is supporting their actions run. But GitHub actions is one of the things that was the most constantly broken without multiple outages recently. So I have the feeling that it proves the point that part of the stability issues is also due to their forced usage of azure.
- bearjaws 1mo agoCentralized source code hosting is going to end up looking like the three credit bureaus in terms of security. It's only a matter of time before the first big hack, when everyone shrugs and says, "Oh well, everyone's source code leaked lol too big to fail."
- smt88 1mo agoGiven the jailbreaking behavior of models in Anthropic and OpenAI, as well as Microsoft allowing Copilot access to client data, we should assume our source code on Github is leaked or leakable anyway.
- hnburnsy 1mo ago>We have since added more than 3 million CPU cores, 120 petabytes of high-speed storage, and significant network capacity. We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure. Crazy.
- ramon156 1mo agothey're still migrating to Azure?
- burstlimit 1mo agoAfter touting 1 billion commits over 2025 at universe last year… they are now handling 3 billion per month jeez. I’ll give them a little more grace after all…
- silver92bullet 1mo agoThis article seems to say that there is "no excuse" for these issues but look at all these things we changed and are changing. It doesn't really feel transparent and it feels like they aren't really taking true ownership on what has happened.
- Preston67 1mo agoThat is amazing
- sleepybrett 1mo agogithub controls the productivity of a large number of very large tech companies. When there is an outage like this they are basically shutting down a significant number of factories for hours at a time. This would be as if during the hayday of detroit they just turned off the power grid at a time randomly at least once a week for hours. It's unacceptable. The amount of productivity lost is staggering. We should be building tools that help us all move off of github as soon as possible. The amount of action code that will need to be rewritten is daunting.
- kburman 1mo agoJust add a queue. Now the outage is eventually consistent. /s
- afgrant 1mo ago“Required several coordinated actions” is the key moment for reflection.
- vladsiu 1mo ago[dead]
- alex7o 1mo agoCant they just put ai agents at optimizing their slow internal paths.
- globular-toast 1mo agoThis seems really bad, to be honest. I think we may be fucked. There's just no way all these lines of code are doing anything useful. We're now just burning stuff in desperation and confusion.
- steve1977 1mo agoMaybe Github (and especially things like Actions) just need to become more expensive?
- _fzslm 1mo agoI appreciate the unprecedented load GitHub is currently experiencing, but it's not just the (admittedly extreme) load of commits/pushes that is to blame. Their Copilot cloud agent offering is suffering with a case of some of the worst corporate ADHD I've seen. We built a cloud agentic development pipeline on it, and it seems like almost every other week they silently change something with zero public announcement that creates real disruption for our team. Note: that's not bugs in the Copilot platform like the article discusses. That's real, breaking changes to the platform that clearly aren't being tested/reviewed before being pushed to prod, with zero public announcement or documentation. Support is useless – we're paying customers in the 4-5 figures and our tickets go unanswered. I love(d) GitHub, but I do think they've lost enough public trust at this point that their time is ticking. With talk of new VCSes designed specifically for agents, I do believe it is just a matter of time. Which pains me somewhat to say.
- hbcdbff 1mo ago1.5 billion commits of worthless slop
- iot_devs 1mo ago> Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. I operated services at similar scale, and generally we use to put a bit of slack so that you would get an alarm when capacity goes up to 80%+ (or whatever number makes sense) This allows to check, in the morning, after coffee, why the load balancer fleet didn't scale up automatically. I am sure there is a good answer to why this is impractical, but it would be nice to know
- afc 1mo ago> Both incidents were capacity failures at their core. We failed to scale critical components before demand exceeded their capacity. This is the wrong way to think about this because there's no such thing as infinite capacity. A large distributed system will be simultaneously mostly idle and (in some subcomponents) overloaded. The root cause is not "a component didn't have enough capacity (because of auto scaling failures)", but rather "this complex system collapses (rather than degrade gracefully) when demand exceeds capacity". When components reach capacity limits, the excess traffic of the lowest priority should be rejected. Rejected traffic should not be retried — in fact, not only should clients not retry these errors, these errors should cause client-side throttling. Traffic isolation should be applied — if the cause of the overload is a single client/customer system, no other system should be affected. Nearly a decade ago I wrote about some of the techniques we applied at Google to implement these protections: https://sre.google/sre-book/handling-overload/ https://sre.google/sre-book/handling-overload/ Most other large internet services have since copied them, afaik.
- eckesicle 1mo agoThis is an excellent book and its lessons saved my bacon many times! As it happens I have a hardcopy of this book (along with "Seeking SRE" and the "SRE Workbook") that I am giving away (because of a move). If you want a hardcopy then email me your UK address I will be happy to post them to your for free. I tried putting them on the street in a little box but surprisingly none of my neighbours grabbed any of my software books. :) EDIT: The books have been given away
- solatic 1mo agoMy last three employers refused to take advantage of Kubernetes PriorityClasses and agree to schedule work to agree (as a cluster-wide resource that affected many teams) on what our PriorityClasses should be and to migrate workloads to have priorities. And this is something relatively easy to implement - no developer work required, and practically no YAML to write. Why not? Because sadly, fundamentally, most workplaces are not run by people who care about day-2 operations or long-term health. Product or Sales pushes customer-visible work into the pipeline, and you dare not say no. "Day-2" work is not considered to be something that moves the needle. Even now, with GitHub facing these severe outages, it's not like they're facing some massive exodus; their load seems to be getting worse over time, not better. I'd be very surprised if there weren't any employees at GitHub who had read the SRE book. I'd expect that they're just not listened to.
- teiferer 1mo ago> Both incidents were capacity failures at their core. We failed to scale critical components before demand exceeded their capacity. I'm missing in these descriptions the most obvious approach: Resilience. Shedding load so that you can keep services up even though capacity is too low. If you flip over as soon as load exceeds what you can handle then this problem will never go away, unless you always have insane overprovisioning of resources which is uneconomical. There will always be spikes. You need to plan to handle them, no matter how high. > we have focused on three priorities: adding capacity, improving efficiency, and removing architectural bottlenecks. Sorry, but again, that is not good enough. They should ask themselves why they are expecting that trying the same medicine as last time will prevent next time. It won't. With that mindset I'm not surprised this happened and it will surely happen again. Edit: In more concrete terms. If you 2x your capacity and in a week you face a burst 2x of what happened last time, you are back in the same seat. If you improve efficiency by 2x, same thing. And after a bottleneck is before a bottleneck. There will always be a bottleneck. The key is to be able to handle a bottleneck. Removing one just pushes the issue to the next one. Your architecture must be such that your whole system should be able to run on a raspberry pi. Most client requests would be dropped, but those that make it through will be served. If your architecture serves 0% because it crashes when load is 10% over capacity, then capacity increases or efficiency increases or bottleneck removal are not going to prevent the next outage.
- deleted 1mo ago[deleted]
- smgpie 1mo agoI have setup a gitea instance on my gitea server which I think is good for me and GitHub both. For one, I dont have to worry about GitHub service outages, and GitHub gets to be free from my toy (and mostly AI slop) projects that no one else will ever read/use/participate in :-)
- mark89h 1mo agonice
- dowonseo 1mo agoNot again but.. the increase in traffic over the last few years is way bigger than I thought
- swedishuser 1mo agoI wonder how much of the traffic increase is enterprise vs. hobbyists? A 7 hour outage for enterprise customers is really, really bad and it's sad if caused by a mass of non-paying vibe coders. It's becoming absolutely obvious that the unlimited free tier needs to go.
- jjice 1mo agoThis is my suspicion, all though I have no evidence, just anecdotes. I make a handful of commits a day. I write code, review it, test it, commit, and then push. We have some marketing folks that have gotten into vibe coding stuff for their personal use. First let me say: good for them and I'm glad they're experimenting with new ideas and tools. The side effect of that is that looking at their repos, they're having Claude go whole hog and make upwards of hundreds of commits a day, all with things that they haven't taken a look at. I don't think I can say this is wrong of them, because their tools encourage that and they shouldn't have to consider their impact on an enterprise service, but I wonder if this trend is similar in other places.
- haul_up 1mo agoThe distinction between running out of capacity and collapsing when you run out of capacity is exactly right. Every distributed system hits limits, the question is what happens next.
- firtoz 1mo agoA lot of these projects and commits would benefit a ton from proper decentralisation. What functionality of GitHub are you *actually* using?
- bob1029 1mo agoI wonder what the ratio of repositories to physical machines is these days. I'd also be curious to see this as change over time. I have a hard time with the premise that a mere doubling of git ops would be especially crippling for any particular repository. GitHub runs like ass because it's oversubscribed by a huge factor. Not because git is inherently constraining at scale.
- drcongo 1mo ago> We have made progress, but these incidents make clear that we must accelerate this work Pretty sure this line appears in every one of these.
- jzer0cool 1mo agoInterview question. How would you handle the growing traffic needs and traffic spikes. I'm curious whether any existing architectural diagram of theirs would highlight a potential failure post-mordem.
- owebmaster 1mo agoIs it a interview question to the MBAs cutting costs at all cost?
- _hzw 1mo agoI recently received a PR fully automated by Claude for an 8 years old repo. The bug is legit and the scope it affects is larger than what that PR addressed, but I no longer care too much about that legacy code anyway, so I also let Claude run free for the first time in my life, from handling that PR to fixing all related bugs. I walked away for half an hour and back, found Claude opened and merged 9 more PRs and added a comprehensive CI for testing for all platforms. It will likely take me months to reach this level of output, but only half an hour for a capable agent. No wonder why GitHub is down all the time.
- frumiousirc 1mo agoLinux didn't (yet) kill Microsoft. Microsoft absorbed that shot. Then the Git arrow went straight to cold black heart of Microsoft. The next few months will determine if they survive it. If they do, what will we see from the third draw out of Linus' quiver?
- inigyou 1mo agoLinux had nothing to do with the downfall of Microsoft. Any perceived dominance of Linux is not because Linux is getting better, but because Microsoft is getting worse.
- prennert 1mo agoWhy does Github not segregate the free offerings from the enterprise or even better, all paid offerings? It is unacceptable that enterprise plans get impacted by traffic on free and public repos. Our repos are neither on the free plan nor are they open. We have not had more AI stuff happening in the last weeks. Our traffic is stable. I would wager that most enterprises did not spike the traffic all of the sudden. Even if they were, we are paying for our quotas. Still our Github actions were breaking and our PRs not viewable at some times. I am hoping this instability is going to cause a Cambrian explosion of forges and if that is happening, Github will be the first victim of the AI revolution. I am working on a truly decentralized / local first code review right now, and a big part of my motivation for this is how bad Github has become. I dont know if I have enough time to build CI as well, but I am hoping others do. Otherwise I will just fall back onto Jenkins.
- kilroy123 1mo agoI agree. I'm just a pro user and not a big enterprise. But I feel I should get a lot more GitHub Actions minutes and stability than free users.
- cs1996 1mo ago"The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR " - silly question but github uses github for their own PRs and deploy right? Do they have a special dedicated system just for them so they can fix github with a code change even if the rest of us can't?
- jjice 1mo agoNot sure, but they could use the self-hosted enterprise version of GitHub for their use case.
- Miyamura80 1mo agoAfaik the other main issue with github downtime is the choice they made a while back coming back to bite them
- mous_tik 1mo agoIs Azure the right choice?
- promptsphere 1mo ago[flagged]
- krupan 1mo agoWhy are we even committing AI generated code and uploading it to GitHub? I've heard we don't need to read the code anymore. Why preserve a detailed version history if AI has it all handled? Why even share code if we can all just have AI write whatever software we need for ourselves? GitHub feels like a dead end of AI is really headed to where we believe it's headed
- lionkor 1mo agoBut GitHub were one of the first and strongest to advocate for using AI, so excessively in fact that it turned people away from the platform. At every turn on the website it suggested to use Copilot, or edit it right now, and commit right now, and do this and that right now(!). The platform effectively begs you to use the free GitHub Actions, too. Not sure what the business model is here. There are a lot of ways to avoid exponentially more commits, issues, and PRs breaking your backend down, and begging every visitor and user to please use AI to write 40x more code that needs 40x more fixes is not one of them.
- augunrik 1mo agoI don't know, this blog post and the lasts read like organisational and human failures to me. Their system doesn't degrade, they even have infinite retries and no meaningful quotas anywhere I can see (but I'm not a heavy github user). I feel like this becomes a lesson on how NOT to design and operate a SaaS.
- plaidfuji 1mo agoThis seems like a pretty straightforward and easily winnable situation for GitHub. The demand for their services just doubled, apparently. They have no real competitor operating at the scale they’re at. They have pretty substantial network effects. They are under no obligation to continue functioning as a bottomless free repository for text file hosting, especially now that text file creation has multiplied exponentially. They could make a few almost purely commercial changes and solve this without any major re-engineering while maintaining their status as the go-to public / open source code hosting platform. 1. Immediately increase pricing of all enterprise licenses and add super-committer overage fees. 2. Rate limit or cap commit size / frequency for public accounts. Their service is more valuable than ever and switching is much harder if people have automation built up on their platform. Now is the time to cash in their chips. And the positive externality of increasing commit cost would be forcing people to have some semblance of restraint for the AI content they generate.
- mitrii 1mo ago[flagged]