4 ms·
> Both incidents were capacity failures at their core. We failed to scale critical components before demand exceeded their capacity. This is the wrong way to t
by afc 1mo ago
> Both incidents were capacity failures at their core. We failed to scale critical components before demand exceeded their capacity.
This is the wrong way to think about this because there's no such thing as infinite capacity. A large distributed system will be simultaneously mostly idle and (in some subcomponents) overloaded. The root cause is not "a component didn't have enough capacity (because of auto scaling failures)", but rather "this complex system collapses (rather than degrade gracefully) when demand exceeds capacity".
When components reach capacity limits, the excess traffic of the lowest priority should be rejected. Rejected traffic should not be retried — in fact, not only should clients not retry these errors, these errors should cause client-side throttling. Traffic isolation should be applied — if the cause of the overload is a single client/customer system, no other system should be affected.
Nearly a decade ago I wrote about some of the techniques we applied at Google to implement these protections: https://sre.google/sre-book/handling-overload/ https://sre.google/sre-book/handling-overload/ Most other large internet services have since copied them, afaik.
- eckesicle 1mo agoThis is an excellent book and its lessons saved my bacon many times! As it happens I have a hardcopy of this book (along with "Seeking SRE" and the "SRE Workbook") that I am giving away (because of a move). If you want a hardcopy then email me your UK address I will be happy to post them to your for free. I tried putting them on the street in a little box but surprisingly none of my neighbours grabbed any of my software books. :) EDIT: The books have been given away
- solatic 1mo agoMy last three employers refused to take advantage of Kubernetes PriorityClasses and agree to schedule work to agree (as a cluster-wide resource that affected many teams) on what our PriorityClasses should be and to migrate workloads to have priorities. And this is something relatively easy to implement - no developer work required, and practically no YAML to write. Why not? Because sadly, fundamentally, most workplaces are not run by people who care about day-2 operations or long-term health. Product or Sales pushes customer-visible work into the pipeline, and you dare not say no. "Day-2" work is not considered to be something that moves the needle. Even now, with GitHub facing these severe outages, it's not like they're facing some massive exodus; their load seems to be getting worse over time, not better. I'd be very surprised if there weren't any employees at GitHub who had read the SRE book. I'd expect that they're just not listened to.
- dilyevsky 1mo agoWhat TP is talking about has nothing to do with workload preemption and is more of a variation of loadshedding (e.g overload management in Envoy). When I was part of the team that ran Google's clusters we had very few priority classes - basically just one for system and majority of serving workload ran on another priority and the rest was for batch. Pretty sure SRE book recommends just that.
- solatic 1mo agoWorkload pre-emption is a form of load-shedding - you shed the load of lower-priority workloads (by evicting their Pods) to free up capacity to schedule more Pods of higher-priority workloads that were added by the Horizontal Pod Autoscaler. > basically just one for system and majority of serving workload ran on another priority and the rest was for batch RCA blames in-house load-balancing services (HAProxy) that reached capacity limits. Even if autoscaling is not working correctly because it didn't take Istio into account - why does it take more than seven hours to just raise the minimum on the Autoscaler for HAProxy and let the workload scheduler evict workloads that are less important than, say, their auth gateway?
- dilyevsky 1mo agoYes, just absolutely crazy way of doing load shedding. > why does it take more than seven hours to just raise the minimum on the Autoscaler for HAProxy and let the workload scheduler evict workloads that are less important than, say, their auth gateway? Like what workloads? Application backends and databases? Did you ever think that your past three employers maybe had a valid point?
- solatic 1mo agoThe entire GitHub site was unavailable. The "unicorn" page. Total outage. Visible to every user. Worst-case scenario. > Application backends? I don't think I'm taking crazy pills to suggest that it's preferable for services like rendering PR diffs, MR merge trains, even accepting new Git commit pushes, to be temporarily unavailable, so that the entire web application doesn't fall over, and cache-friendly read-only workloads continue to succeed.
- vasco 1mo agoOf course there's infinite capacity. A google datacenter is infinite capacity from the perspective of say an NTP server. Infinites exist when you have enough orders of magnitude in the middle. Also traffic isolation and degradation by tier is not "no outage", you're still in outage land, you're just being smart in how you use it and choosing what you disrupt. It doesn't fix the lack of capacity.
- nosefrog 1mo agoI don't understand your comment. A google data center is much larger than an ntp server, but it's obviously not infinitely larger. As you know, if it was infinite capacity, then there would be no need for load balancing or load shedding. And of course, load shedding low priority traffic is still a partial outage, it's just a less bad outage than load shedding high priority traffic. It does not fix lack of capacity, but it significantly lessens the negative effects of it.
- Brian_K_White 1mo agoIt is infinitely larger, because there is no distinction between more than you use, and infinity.
- catlifeonmars 1mo agoSure there is infinite capacity given an infinite amount of time to scale up. I assert you’re leaving out the time dimension. Those so-called infinities are simply not accessible in a practical way since you’ll hit a wall in actually provisioning that capacity long before the data center runs out of compute.
- vasco 1mo agoYou understand it if you think of engineering infinites rather than mathematical infinites. The capacity of a full single datacenter can be treated as infinite for most customers. I explained how its defined in the original comment. Amount of places in engineering where you treat even a 3 order of magnitude difference as infinite is a lot, but the number of order of magnitudes varies depending on context.
- d0vs 1mo agoGH didn't collapse so it's not impossible that they already implement these measures: > At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. https://www.githubstatus.com/incidents/zkxwbgr0cnmx https://www.githubstatus.com/incidents/zkxwbgr0cnmx
- xnorswap 1mo agoIt completely collapsed, 20% is meaningless, as I said in: https://news.ycombinator.com/item?id=49333089 https://news.ycombinator.com/item?id=49333089 If you load an issue page, you'll see 1 failed request to: /project/product/issues/<number> And sure, that's what you care about, but consider the working requests to: /in-product-messaging/copilot-budget-request-banner /in-product-messaging/code-scanning-ai-findings-preview-banner /github-copilot/chat /_private/browser/stats Those are actual endpoints and results.
- concerned_user 1mo agoIf page has 15 requests and needs data from all of them to work correctly, then with 20% failure rate you are suddenly close to 100% non-functional page from the user perspective.
- afc 1mo ago+1. This is a very real problem in practice. A technique we've used to deal with this situation: in the overloaded backend (that has to reject some percentage of incoming requests), group the incoming requests by the parent request (the one with the 1:15 fan out) and reject according to the parent request. One way to put it, simply (though somewhat inaccurately), would be: reject 100% of traffic from 20% of users, rather than 20% of traffic across all users (causing essentially full failure for all users). We typically implemented this by propagating an ID of the parent request down to the backend. I'm simplifying a lot in this description (e.g. have to deal with the parent requests landing on different backend tasks; also rotate the IDs gradually to introduce some fairness).
- neya 1mo agoLet's be honest here, the real reason is the crap that is Azure. GitHub was perfectly fine until then. They are just too bureaucratic to admit it. Worth reading: https://isolveproblems.substack.com/p/how-microsoft-vaporized-a-trillion https://isolveproblems.substack.com/p/how-microsoft-vaporize...
- TiredOfLife 1mo agoSo for the past 6 years the bad uptime was because of less than 12% of github being on azure?
- XzAeRosho 1mo agoRead the article. It's about Microsoft culture messing things up, not Azure themselves (though it may be a factor).
- Medowar 1mo agoactually, I think this time it is the other way around, Azure is more the solution than the problem. Github in the past ran on their own Hardware. That is fine, if your load is predictable and nto changing rapidly. However, the evolution of the past few months/years has shown, that the previous assumptions about growth are now outdated and scaling that capacity on your own metal is not that easy. Hardware has lead times of many weeks, especially in the current situation, datacenter capacity is even longer and more difficult, especially right now. Choosing not to deal with scaling the hardware is a valid choice in this situation. Yes, Azure is a bunch of servers held together with glue, duct tape and a lot of hope, but I think, the github hardware is not much better at the moment.
- neya 1mo agoAgreed on not hosting things yourself, but, I am not arguing against cloud hosting at all - just Azure. If they just came out and admitted it is a disaster and swallowed their pride and moved their services towards anything else at all - GCP, AWS or whatever else - I think their uptime would significantly improve. Of course, the fundamental problem here is the culture of the company itself. That's harder to fix.
- ChrisMarshallNY 1mo ago> this complex system collapses (rather than degrade gracefully) when demand exceeds capacity This. It's easy to be an "armchair quarterback," here, though. Handling stuff like this, needs to be planned for, from the start. I suspect that a lot of the issues are because GitHub is something that started small (and probably quickly), and has accreted. Things like Facebook are in a similar boat.
- afc 1mo agoHaha, tell me about it! I spent more than a decade mostly just deploying these systems across just one company. It's far from trivial!
- yard2010 1mo agoThank you so much for this gem, this sent me through the rabbit hole and I'm fascinated by the level of engineering in this book. I love how big complex systems are engineered. It feels like anything is possible when you have a solid plan and you keep improving it step by step. So inspiring.
- 13639366668 1mo ago[flagged]
- abustamam 1mo agoIsn't this exactly the kind of systems design questions they ask during interviews??? Makes me wonder how many folks at GH and MSFT could even pass their own interviews. Thanks for sharing though. I learn more about systems design from HN comments than anything else