16 ms·
Lessons Learned from Twenty Years of Site Reliability Engineering
- deleted 3y ago[deleted]
- 6LLvveMx2koXfwn 3y ago"We once narrowly missed a major outage because the engineer who submitted the would-be-triggering change unplugged their desktop computer before the change could propagate". Sorry, what?
- jrms 3y agoI thought the same.
- jedberg 3y agoThe change was being orchestrated from their desktop, and they noticed thing were going sideways, so they unplugged their desktop to stop the deployment. Aka "pressed the big red button".
- francisofascii 3y agoYeah, interesting tidbit. It might sound insane today that one engineer's desktop computer could cause such an outage. But that was probably more commonplace 20 years ago and even today in smaller orgs.
- shadowgovt 3y agoThere was a famous incident at one point where code search internally went down. It turned out that while they had deployed the tool internally, one piece of the indexing process was still running as a cron job on the original developer's desktop machine. He went on vacation, his credentials aged out, and of crawler stopped refreshing the index. But my favorite incident will forever be the time they had to drill out a safe because they were disaster-testing the password vault system and discovered that the key needed to restore the password vault system was stored in s aafe, the combination for which had been moved into the password vault system. Only with advanced, modern technology can you lock the keys to the safe in the safe itself with so many steps!
- tomcam 3y ago> But my favorite incident will forever be the time they had to drill out a safe because they were disaster-testing the password vault system Great story to be sure—-but I’m going to call it a success. They did the end-to-end testing and caught it then instead of real life.
- shadowgovt 3y agoWell, mostly-kinda-sorta. ;) It's the internal password vault, and there's only one of them, so it's more like "they broke it on purpose and then had to fix it before the company went off the rails." Among the things kept in that vault are credentials that if they age out or aren't regularly refreshed, key internal infrastructure starts grinding to a halt. But still, "it broke while engineers were staring at it and trying to break it" is a better scenario than "it broke surprisingly while engineers were trying to do something else."
- dilyevsky 3y agoAt one point I had to run a script on a substantial portion of their server fleet (like hundreds of thousands machines) and I remember I ran it with a pssh-style utility from desktop (was 10y ago so dunno if they still use this). It was surprisingly quick to do it this way. Could’ve been something like that
- codemac 3y agoIt always cracks me up how Google is simultaneously the most web-based company in the world, but their internal political landscape was (Infra, Search, Ads) > everything else. This leads to infra swe writing stupid CLIs all day, rather than having any literal buttons. Things were changing a lot by the time I left though. I do think Google should be more open about their internal outages. This one in particular was very famous internally.
- dilyevsky 3y agoWe also avoided some outages by running one-off scripts fleet-wide so it cuts both ways
- throwawaaarrgh 3y agoEvery large enterprise is internally a tire fire.
- jeffbee 3y agoIt's a logical consequence of the "zero trust" network. If an engineer's workstation can make RPCs to production systems, and that engineer is properly entitled to assume some privileged role, then there's no difference between running the automation in prod and running it on your workstation. Even at huge scales, shell tools plus RPC client CLIs can contact every machine in the world pretty promptly.
- kevan 3y agoThere's still differences. If you're running it in prod then the functionality has at least gone through code review and you have higher confidence what's running is what you think it is. If you run things from personal boxes there's always the risk of them not having the latest code, having made a local change and not checking it in, or the worst case of a bad actor doing whatever they want with the privileged role. But if code review isn't required or engineers have unrestricted SSH access to prod hosts then it's pretty much equivalent.
- NiloCK 3y agoTo think - if it'd been a laptop they would have had to smash it with a hammer.
- jedberg 3y agoThis is a great writeup, and very broadly applicable. I don't see any "this would only apply at Google" in here. > COMMUNICATION CHANNELS! AND BACKUP CHANNELS!! AND BACKUPS FOR THOSE BACKUP CHANNELS!!! Yes! At Netflix, when we picked vendors for systems that we used during an outage, we always had to make sure they were not on AWS. At reddit we had a server in the office with a backup IRC server in case the main one we used was unavailable. And I think Google has a backup IRC server on AWS, but that might just apocryphal. Always good to make sure you have a backup side channel that has as little to do with your infrastructure as possible.
- dilyevsky 3y agoAfair Google just ran irc on their corp network which was completely separate from prod so I wouldn’t be surprised if it was in a small server room in the office somewhere. > very broadly applicable. I don't see any "this would only apply at Google" in here. One thing I haven’t even heard of anyone else doing was production panic rooms - a secure room with backup vpn to prod
- DaiPlusPlus 3y ago> Afair Google just ran irc on their corp network which was completely separate from prod I thought Google didn't have a "corp" network because of their embrace of zero-trust in BeyondCorp?
- dilyevsky 3y agoI don't think zero-trust prohibits network segmentation for redundancy or due to geographical constraints etc. It's mainly about how you gain access.
- Moto7451 3y agoCorrect. At an old job we did zero trust corp on a different AWS region and account. The admin site was a different zero trust zone in prod region/account and was supposed to eventually become another AWS account in another region (for cost purposes). I can’t say if any of this was ideal but it did work unobtrusively.
- tiddo_langerak 3y agoI'm curious how people approach big red buttons and intentional graceful degradation in practice, and especially how to ensure that these work when the system is experiencing problems. E.g. do you use db-based "feature flags"? What do you do then if the DB itself is overloaded, or the API through which you access the DB? Or do you use static startup flags (e.g. env variables)? How do you ensure these get rolled out quickly enough? Something else entirely?
- shadowgovt 3y agoWhen you're a small company, simpler is actually better... It's best to keep it simple so that recovery is easy over building out a more complicated solution that is reliable in the average case but fragile in the limits. Even if that means there's some places on the critical path where you don't use double redundancy but as a result the system is simple enough to fit in the heads of all the maintainers and can be rebooted or reverted easily. ... But once your firm starts making guarantees like "five nines uptime," there will be some complexity necessary to devise a system that can continue to be developed and improved while maintaining those guarantees.
- dilyevsky 3y agoThere’s a chapter on client-side throttling in the sre book - https://sre.google/sre-book/handling-overload/ https://sre.google/sre-book/handling-overload/ At google we also had to routinely do “backend drains” of particular clusters when we deemed them unhealthy and they had a system to do that quickly at the api/lb layer. At other places I’ve also seen that done with application level flags so you’d do kubectl edit which is obviously less than ideal but worked
- justapassenger 3y agoImplantation details will depend on your stack, but 3 main things I’d keep in mind: 1. Keep it simple. No elaborate logic. No complex data stores. Just a simple checking of the flag. 2. Do it as close to the source as possible, but have limited trust in your clients - you may have old versions, things not propagating, bugs, etc. So best to have option to degrade both in the client and on the server. If you can do only one, so the server side. 3. Real world test! And test often. Don’t trust test environment. Test on real world traffic. Do periodic tests at small scale (like 0.1% of traffic) but also do more full scale tests on a schedule. If you didn’t test it, it won’t work when you need it. If it worked a year ago, it will likely not work now. If it’s not tested, it’ll likely cause more damage than it’ll solve.
- Dowwie 3y agoIf you're using AWS resources, give LocalStack a try for integration testing
- whummer 3y ago100% - Including some of the Chaos Engineering features that are recently offered in the platform (e.g., simulating service errors/latencies, region outages, etc)
- galkk 3y agoWife used localstack at previous job and it was miserable experience. Especially their emulation of queues. Maybe things have improved since couple years ago, though
- fensterblick 3y agoRecently, I've heard of several companies folding up SRE and moving individuals to their SWE teams. Rumors are that LinkedIn, Adobe, and Robinhood have done this. This made me think: is SRE a byproduct of a bubble economy of easy money? Why not operate without the significant added expense of SRE teams? I wonder if SRE will be around 10 years from now, much like Sys Admins and QA testers have mostly disappeared. Instead, many of those functions are performed by software development teams.
- gen220 3y agoI imagine the threshold is something like 1 SRE for every $1mm of high-margin revenue you can link to guaranteeing the 2nd "9" of $product availability/reliability.
- donalhunt 3y agoI believe that is indeed a good guide for when it makes sense to have a SRE team supporting a service or product (with the caveat that the number probably isn't $1MM). There are also good patterns for ensuring you actually have adequate SRE coverage for what the business needs. 2 x 6ppl teams geo-graphically dispersed doing 7x12 shifts works pretty well (not cheap). You can do it with less but you run into more challenges when individuals leave / get burnt out / etc.
- oceanplexian 3y agoThat’s sort of ridiculous. A mid-level SRE easily costs a quarter of that. And a company like Apple would then have 80,000 SREs? Lol no.
- gen220 3y agoI think you've perhaps misread my post? It's marginal revenue attributable to a high-performing SRE (i.e. an SRE who would be able to elevate a product they're supporting from 90.0% availability to 99.0% availability. It's actually a pretty high bar, because there aren't that many products for which the that segment of availability translates to >$1mm in marginal revenue. $1mm is a ballpark figure, but I think it's the right order of magnitude (i.e. the true number might be $5mm). Expanding on another point in the original post: decision varies with the profitability of that marginal revenue. For example, it's basically pure profit for Google, Amazon or Netflix – accordingly, it makes sense that they'd have many people who focus exclusively on performance and availability, to make sure they aren't leaving that revenue on the ground.
- Smaug123 3y agoFor much much more on this, I'm most of the way through Google's book _Building Secure and Reliable Systems_, which is a proper textbook (not light reading). It's a pretty interesting book! A lot of what it says is just common sense, but as the saying goes, "common sense" is an oxymoron; it's felt useful to have refreshed my knowledge of the whole thing at once.
- tap-snap-or-nap 3y agoFor those who want to read it https://google.github.io/building-secure-and-reliable-systems/raw/toc.html https://google.github.io/building-secure-and-reliable-system...
- benlivengood 3y agoSomething I hope to eventually hear is the solution to the full cold start problem. Most giant custom-stack companies have circular dependencies on core infrastructure. Software-defined networking needs some software running to start routing packets again, diskless machines need some storage to boot from, authentication services need access to storage to start handing out service credentials to bootstrap secure authz, etc. It's currently handled by running many independent regions so that data centers can be brought up from fully dark by bootstrapping them from existing infra. I haven't heard of anyone bringing the stack up from a full power-off situation. Even when Facebook completely broke its production network a couple years ago the machines stayed on and running and had some internal connectivity. This matters to everyone because while cloud resources are great at automatic restarts and fault recovery there's no guarantee that AWS, GCP, and friends would come back up after, e.g., a massive solar storm that knocks out the grid worldwide for long enough to run the generators down. My guess is that there are some dedicated small DCs with exceptional backup power and the ability to be fully isolated from grid surges (flywheel transformers or similar).
- Gh0stRAT 3y agoAzure has procedures in place to prevent circular dependencies, and regularly exercises them when bringing new regions online. IIRC some of the information about their approach is considered sensitive so I won't elaborate further.
- jeremyjh 3y agoAre you saying they can bring a new data center online without any connectivity to the rest of their infrastructure? GP isn't concerned about turning on one data center, they are concerned about turning them all on at the same time, and that can never be tested.
- Gh0stRAT 3y agoIf you practice bringing a new datacenter online without any connectivity on the existing deployment, and you practice then joining two disjoint "clouds", then you've pretty much covered your bases. Are you making a rate limiting/ddos argument about "turning them all on at the same time"?
- alexpotato 3y agoIf you are interested in a similar list but with a bent towards being a SRE for 15 years in FinTech/Banks/Hedge Funds/Crypto, let me humbly suggest you check out: https://x.com/alexpotato/status/1432302823383998471?s=20 https://x.com/alexpotato/status/1432302823383998471?s=20 Teaser: "25. If you have a rules engine where it's easier to make a new rule than to find an existing rule based on filter criteria: you will end up with lots of duplicate rules."
- xyst 3y agoOff topic: TIL Google has its own TLD (.google)
- DaiPlusPlus 3y agoSo does .airbus, .barclays, ,mcrosoft, and .travelersinsurance too - it's nothing new https://data.iana.org/TLD/tlds-alpha-by-domain.txt https://data.iana.org/TLD/tlds-alpha-by-domain.txt
- Racing0461 3y agoFor a sr sde yearly salary, you can own one too. The application process is nevertheless, "may issue".
- toast0 3y agoAre they accepting applications again, or do you also need a time machine?
- teddyh 3y agoFrom what I can tell, Google owns at least eleven TLDs, just for themselves: • .android • .cal • .chrome • .gbiz • .gle • .gmail • .goog • .google • .play • .prod • .youtube Google also owns 22 generic domains: • .app • .boo • .channel • .dad • .day • .dev • .eat • .esq • .fly • .foo • .hangout • .here • .how • .ing • .meme • .mov • .new • .nexus • .page • .prof • .search • .zip
- sumedh 3y agoHow did Apple allow Google to take over .app
- throwawaaarrgh 3y agoThe cheapest way to prevent an outage is to catch it early in its lifecycle. Software bugs are like real bugs. First is the egg, that's the idea of the change. Then there's the nymph, when it hatches; first POC. By the time it hits production, it's an adult. Wait - isn't there a stage before adulthood? Yes! Your app should have several stages of maturity before it reaches adulthood. It's far cheaper to find that bug before it becomes fully grown (and starts laying its own eggs!) If you can't do canaries and rollbacks are problematic, add more testing before the production deploy. Linters, unit tests, end to end tests, profilers, synthetic monitors, read-only copies of production, performance tests, etc. Use as many ways as you can to find the bug early. Feature flags, backwards compatibility, and other methods are also useful. But nothing beats Shift Left.
- 8040 3y agoI would like to take this moment to really highlight "Recovery mechanisms should be fully tested before an emergency". As the unexpected SRE at Google who became known by entire company for using a double negative incorrectly, it is something very important to do right away. For those Googlers curious, you can search my username internally for how I generated more impact then could be measured.
- _boffin_ 3y agoPossible to give more insightful details?
- js2 3y ago> Automate your mitigations Think long and hard about this one. Multiple times in my three-decade career I've seen automated mitigations make the problem worst. So really consider whether self-healing is something you need. I built my company's in-house mobile crash reporting solution in 2014. Part of the backend has had one server running Redis as a single point of failure. The failover process is only semi-automated. A human has to initiate it after confirming alerts about it being down are valid. There's also no real financial cost to it going down - at worst my company's mobile app developers are inconvenienced for a bit. In the decade the system has been operational I can count on two fingers the number of times I've had to failover. Despite this system having no SLA it's had better uptime than much more critical internal systems. Conversely: https://github.blog/2023-05-16-addressing-githubs-recent-availability-issues/ https://github.blog/2023-05-16-addressing-githubs-recent-ava... https://github.blog/2018-10-30-oct21-post-incident-analysis/ https://github.blog/2018-10-30-oct21-post-incident-analysis/ https://www.datacenterknowledge.com/archives/2012/12/27/github-outage-caused-by-failover-snafu https://www.datacenterknowledge.com/archives/2012/12/27/gith... To be fair, GitHub operates at a much larger scale. My point is only that redundancy and automated mitigations add complexity and are almost by definition rarely tested and operate under unforeseen circumstances. So really, consider your SLA and the cost of an outage and balance that against the complexity you'll add by guarding against an outage. I think my first introduction to this was circa 1998 when I had a pair of NetApps clustered into an HA configuration and one of them failed and caused the other to corrupt all its disks. Fun times. A similar thing happened around the same time with a pair of Cisco PIX firewalls. I've been leery of HA and automated failover/mitigations ever since.
- seedless-sensat 3y agoI find it interesting that this reflection didn't mention SLI/SLO/error budgets, which Google SRE has championed for a long time. My impression is that they're nice in theory, but less useful in practice. I'm yet to see an error budget effectively inform eng decision making.
- jabroni_salad 3y agoError budgets are to control the workload of the guy who is holding the oncall pager, who otherwise has no say over his or her situation. In recent years companies have shifted to 'you build it you own it' and the infra has been abstracted to the point that the SWE can own the entire thing. Error budgets also only matter if you either give a shit about your guys or have to pay them for that oncall time. Plenty of employers are happy so say 'salary is exempt, suck it up lol' so errors are effectively free.
- korginator 3y agoThis feels very apropos to a recent banking outage we had here in Singapore. On Oct 14, I was out buying groceries and was asked a very strange question by the NTUC FairPrice supermarket cashier, "which bank is your card from?". Expecting the usual "would you like this product on offer" type question, I didn't even register the question for a second. Turns out that we had a major outage at a data centre that served DBS - one of the largest lenders, as well as Citi. [1] The disruption was attributed to a cooling system failure at a data centre operated by Equinix [2]. Further digging led to information that the culprit was the SG3 data centre [3] marketed as their largest IBX data centre in the Asia Pacific region and one of the newest. It turned out that the cooling system was being upgraded on contract with an external vendor, who applied incorrect settings which brought it down. Further, this particular data centre has 2N electrical redundancy but a cooling redundancy of only N+1 chillers, in comparison to other financial services organizations like the Singapore exchange (SGX) that offers [4] CoLo hosting with 2N chillers, which I believe is essential for warm equatorial climates. Sadly, this outage was followed by yet another smaller payment related outage the following week, making it the fifth outage this year. DBS was trumpeting their move to the cloud [5] as part of their grand plan to transform themselves from a bank into a software company that also offered banking services [6][7]. In going all out with this questionable and misguided transformation they've lost focus on what made people trust them in the first place - the decades of trust that was built on solid, reliable, transparent and efficient banking services. There were questions about why their backup data centre didn't kick in and there are no answers till date. It's clear to see that the recovery mechanisms weren't tested, performance degrade modes were not implemented or not tested, disaster resilience utterly failed, there were no working mitigations for cooling system failures or DC failures, and as a result ATMs and other services were down from 3pm on Oct 14 until the following morning. For a country that prides itself on digital transformation, this is just the latest banking systems failure which makes it far more than an egg in the face, it's an erosion of trust. Items (2), (3), (7), (8), (9) from the Google report directly apply to this failure. I can only hope something good comes out of it and lessons are learnt. [1] https://www.channelnewsasia.com/singapore/dbs-citibank-outage-data-centre-cooling-system-down-3861076 https://www.channelnewsasia.com/singapore/dbs-citibank-outag... [2] https://www.zdnet.com/article/equinixs-data-center-system-upgrade-results-in-hours-long-disruption-at-banks/ https://www.zdnet.com/article/equinixs-data-center-system-up... [3] https://www.equinix.com/data-centers/asia-pacific-colocation/singapore-colocation/singapore-data-center/sg3 https://www.equinix.com/data-centers/asia-pacific-colocation... [4] https://www.sgx.com/data-connectivity/co-location https://www.sgx.com/data-connectivity/co-location [5] https://www.dbs.com/newsroom/First_bank_in_Singapore_to_launch_new_cloud_based_data_centre https://www.dbs.com/newsroom/First_bank_in_Singapore_to_laun... [6] https://bankinginnovation.qorusglobal.com/content/articles/transforming-dbs-bank-tech-company https://bankinginnovation.qorusglobal.com/content/articles/t... [7] https://www.dbs.com/technology-future/dbs-redefining-the-future-of-banking-as-a-technology-company.html https://www.dbs.com/technology-future/dbs-redefining-the-fut...