7 ms·
Are these outages caused by introduced bugs, though, or by load issues? As someone who has spent many years working in high load environments, this is not an u
by cortesoft 2mo ago
Are these outages caused by introduced bugs, though, or by load issues?
As someone who has spent many years working in high load environments, this is not an uncommon pattern.
You design a system and it works great. It can handle failures, load spikes, it is horizontally scalable, things are great. You think you figured it out.
And then load keeps increasing and you suddenly hit a tipping point where everything keeps failing, and you cant keep up. The things that you thought were perfectly horizontally scalable turn out to have a bottleneck you didn’t even think about until you got to a truly massive scale. Your systems suddenly don’t have the excess capacity to handle load spikes or catchup work, so suddenly any failure cascades and recovery is more and more difficult. You can’t solve the problem with additional hardware, and your perfect scalable design actually can’t scale any more.
This doesn’t have to be about GitHub using LLMs in their code to still be related to LLMs. GitHub gets a lot more commits now because of LLMs and probably get a lot more reads because of LLMs as well.
It could be that the extra usage just pushed them past one of those capacity thresholds.
- Syntaf 2mo agoThere was a great article awhile back that shed some light on just how dysfunctional azure is as a platform: https://news.ycombinator.com/item?id=47616242 https://news.ycombinator.com/item?id=47616242 If I had to guess it's because Github is sitting on top on infrastructure held up by toothpicks and duct tape
- toomuchtodo 2mo agoWhich is somewhat humorous because it was arguably more stable when they ran on their own hosted colo infra before moving to Azure. This was a choice versus keeping the infra compartmentalized and using Azure for elastic overflow compute needs. I'm sure marketing and bonuses rest on throwing it all on the Azure quicksand though. GitHub Will Prioritize Migrating to Azure Over Feature Development - https://news.ycombinator.com/item?id=45517173 https://news.ycombinator.com/item?id=45517173 - October 2025 (63 comments)
- hirako2000 2mo agoIt pleases shareholders more to hear about exponential Cloud offering adoption than SLA availability for a developer platform.
- b112 2mo agoExcept, lots of devs use github, and see what using Azure means. Management is seeing that downtime too. So what insane person will ever touch Azure willingly? Hotmail used to run on bsd, and they tried to move it to WinNT, back in the day. Same thing. Disaster. And it was held up as a reason to never use winnt for serious work, for years. Microsoft just always shoots themselves in the foot.
- vladvasiliu 2mo ago> So what insane person will ever touch Azure willingly? The nobody-got-fired-for-choosing-Microsoft crowd. And they don’t look few and far between.
- porridgeraisin 2mo agoAnd also the most important few thousand accounts will get custom support and all sorts of patchwork and MS will put the effort to solve all their problems. So the customers that matter to MS are likely having a great time on azure. And for every ten thousand of us chums that leave azure, there's probably one sweetheart deal won through wine and dine that more than compensates.
- zamalek 2mo agoI think my age group (~40) is the tail end of widespread dogmatic trust in Microsoft. Once we retire I think they are going to be in serious trouble.
- lenerdenator 2mo agoThey don't need dogmatic trust, just oligopolies and monopolies.
- BedVibe_Studios 2mo agoIf that is truth its scary of how many projects use that cause it's supposed to be reliable.
- mirsadm 2mo agoDunno that's my experience with most businesses. I doubt AWS is that much better under the hood.
- MarkMarine 2mo agoHaven't worked in AWS but as a customer of AWS, GCP, and Azure there is a quite a difference between running workloads that are doing long-running computations on AWS (on demand is so rock solid that loosing a box is a rarity, I run on spot and still complete almost all of my runs) and running on Azure's "on-demand" style offering feels like spot (or worse) on AWS.
- vel0city 2mo ago> on demand is so rock solid that loosing a box is a rarit Maybe this really depends on region, but I can't say this is true. I've experienced tons of hardware failures that have caused instances to throw weird errors, instances to randomly stop, instances to randomly disappear (along with their corresponding EBS volume). Standard EBS volumes are only 99.8% durable. If you run a lot of instances on EBS, some of them will disappear. GCP meanwhile has balanced zonal PDs (not even regional, zonal) at >99.999% durability. I've never lost a disk on GCP, I've never even had an instance stop once without me telling it to stop. Azure is absolutely a hot mess though. Had a VM a client was paying for backups on. Couldn't restore the backup because they changed generations of VM platforms too many times, literally no way to restore it. Had to get their support to eventually give me a .vhd that magically appeared in a OneDrive share a few days after asking about it.
- Nextgrid 2mo ago> Are these outages caused by introduced bugs, though, or by load issues? Probably both. But it still throws a thorn into the theory that LLMs are about to replace software engineers any day now. You'd think they could LLM-code their way out of this situation easily if LLMs were the software engineer replacement they are being marketed as.
- hirako2000 2mo agoBut that's just a perfectly fitting excuse for GitHub degradation. Web search, steaming, high frequency trading, and many other systems are resource demanding, and keep scaling. Whether to handle the surge due to bots or wider adoption. But GitHub can't scale git? It isn't even git failing. Incompetence in leadership is what makes a tech business technically unable to meet growing demand.
- remus 2mo ago> Incompetence in leadership is what makes a tech business technically unable to meet growing demand. I don't think this is true. Take LLMs for example, you need GPUs to serve them and these are in short supply at the moment making it hard to meet demand. I don't think this is necessarily the fault of incompetent leadership.
- inigyou 2mo agoMaybe competent leadership would have spun up a foundry 5 years ago?
- hirako2000 2mo agoThe exception doesn't make the rule.
- ifwinterco 2mo agoGitHub is different because they need to scale writes, that’s fundamentally a much much harder problem
- hirako2000 2mo agoWhich wasn't a problem before MS acquisition. Git is the write intensive process. It scaled to millions of contributors. Then couldn't keep scaling. But sure if you believe the narrative and blame ai bots.
- ModernMech 2mo ago
- glaslong 2mo agoActions is one thing, that seems like a difficult, dynamic and bursty thing to host, even before the Vibe Cambrian Explosion. And _nothing_ at scale is easy... But Pages? The static sites? Down for so many hours? Oof.
- Grimburger 2mo agoThis is an excellent comment and I don't want to diminish by adding jokes but this totally reminds me of the hitler uses kubernetes meme. You can do everything right for HA and there'll still be something to trip you up. https://www.reddit.com/r/kubernetes/comments/f2bcsu/parodyfunny_hitler_uses_kubernetes/ https://www.reddit.com/r/kubernetes/comments/f2bcsu/parodyfu...
- paddybyers 2mo agoI agree, no horizontally-scalable system is infinitely scalable in any dimension. You will hit some limit - for example, you might be able to support arbitrary scale in some dimension, but there is a bound on the sustainable rate of change, or the second-order rate of change, and you hit that limit. You seem to be suggesting though that hitting those limits came as a surprise, and/or there was no future architecture plan to address that when it happens, which I think is the real problem here. You would expect a team with the maturity of Github to understand what those limits are, and plan for them, based on predictable growth in demand.
- strangattractor 2mo agoUse of Claude for coding increases the number of commits, merges and pushes. Add in difficulty in getting RAM, Disk Storage and Servers. A dash of high energy prices, data center angst and tariffs. A pinch of diverting resources to all things AI. You get legacy systems that cannot keep up with demand for expansion.
- SoftTalker 2mo agoOne of the big motivating factors in the development of git was the desire for decentralized version control. With GitHub, we threw that away. It's centralized on steroids. Now we're all depending on one platform that has to be massively scaled to deal with the massive load of serving almost every major software project on the planet. And when it is down, we all notice.
- pythonaut_16 2mo ago> With Github, we threw that away. I see this repeated a lot and IMO it's simply not true. The main decentralized advantage of Git is that you can continue to do VCS operations without network access or access to the remote host. Most of what Github does is managing collaboration. The only thing we've "thrown away" by using Github is the email based workflows or directly pushing git branches to various hosts. But unless you're gonna have team members SSH into each other's machines you'd still need a central repo somewhere.
- talonx 2mo agoAI slop-code has exacerbated the problem and GitHub is not able to keep up. If you look at the last 12 month's outage reports from GitHub's own status page, many of them mention capacity issues as an underlying cause.