2 ms·
Which wasn't a problem before MS acquisition. Git is the write intensive process. It scaled to millions of contributors. Then couldn't keep scaling. But sure if
by hirako2000 2mo ago
Which wasn't a problem before MS acquisition. Git is the write intensive process. It scaled to millions of contributors. Then couldn't keep scaling. But sure if you believe the narrative and blame ai bots.
- ModernMech 2mo agoGitHub had 30 million users in 2018 when it was acquired by Microsoft. This year alone they added 30 million users for total of 180 million. So pre-Microsoft GitHub faced very different problems compared to post-Microsoft GitHub not even considering poor management or AI. They're by far the largest platform of its kind, so I'm sure the problems are uniquely difficult to solve.
- Grombobulous 2mo agoI have trouble believing it’s a uniquely difficult situation when no other tech company of the same or larger scale has those same problems. Is there a single Google service that has ever had reliability this bad? Facebook? Instagram? TikTok? Apple? Those companies all run hugely scaled write-heavy platforms. Why is GitHub so uniquely unreliable? Me and 1 billion of my best friends can upload 5GB 4K videos to iCloud Photos and live stream to Facebook all day long but GitHub who is owned by Microsoft the second largest cloud computing provider on the planet can’t handle 180 million users pushing code and running pipelines on VMs?
- ModernMech 2mo agoI won't pretend to be an expert in the various systems of running these services, but to me it seems that GitHub's system is especially vulnerable to what AI is doing. I'll give you an example: I pointed Codex at one of my repositories asking it to implement some features. Very quickly, it ballooned a 3 minute CI workflow to 30 minutes per commit, and the number of commits it started making increased 10 fold. So we're looking at a 100x increase in usage just from myself. You factor that into the fact that all public repo get access to this same compute, basically unlimited minutes, up to 20 parallel tasks and 6 hours per request. Now because of AI, every small little weekend throw away side project is running a full dev ops shop. And the pressure never goes down because the AI just start a cron job to run these systems every single day, whether anyone is even looking at these repos including the authors, who may themselves just be bots. And the systems are sophisticated: parallel builds across windows, mac, multiple linux distributions, multiple architectures, embedded targets, web targets, full feature matrix, the works. The default MO of these agents is to do as much testing as possible without any regard for resource constraints, and GitHub offers them unlimited resources, so it's like an addict meeting a dealer. So in a sense, maybe being attached to Microsoft is the problem. Or at least their unlimited wallet is -- Microsoft is the enabler in all of this.
- Grombobulous 2mo agoI also can't pretend to be an expert on this. Conceptually, everything you're saying makes sense. It still just sounds like a really vanilla "horizontally scaling VMs" problem, especially for yesterday's incident that was focused on pipelines. If this is a "GitHub is giving away more capacity than it has" problem, that's easily solved by rate limiting and queuing. This is GitHub's explanation: > During the incident, some Actions Runner Controller (ARC) runner pods became stuck in an idle state. Affected users can delete those pods using kubectl or redeploy their Actions Runner Controller application. ARC will automatically create replacement runners. > The next releases of Actions Runner and Actions Runner Controller will include an automatic recovery mechanism, preventing the need for these manual steps in the future. That sounds a lot more like a major architectural flaw and a really obvious oversight.