6 ms·
Operating a large distributed system in a reliable way: practices I learned
- joshgel 7y ago> I like to think of the effort to operate a distributed system being similar to operating a large organization, like a hospital. Clearly never worked for a hospital. Hospitals need good engineers (and often don’t have them). Our ‘nines’ are embarrassing...
- jpitz 7y agoAre you referring to medical operations, or IT operations, in hospitals? I think he is referring to medical operations, where I would expect the relevant professions to be doctors and nurses, not engineers.
- arkades 7y agoActually large hospital systems tend to hire one or two systems engineers to be part of the QI department. But yeah, most QI is front line staff.
- ambicapter 7y agoQI department?
- arkades 7y agoQuality improvement.
- ikiris 7y agoThe point stands. I've been involved in big hospital management and at a FAANG. Hospitals are a horror show if you see behind the curtain in both respects, and others.
- joshgel 7y agoMedicine works not because we have learned to scale it in any sense, but on the contrary, because we still rely on individual physicians and nurses to provide care. So like the HR system of large hospitals is probably better/bigger than small ones in terms of on boardinging doctors, but the care at a large center is only as good as the individual physician or nurse caring for patients. Probably bigger, academic centers are (very slightly) better (though this is an active debate with recent literature in high profile journals on both sides), but if they are, I suspect its because they can recruit better individual doctors. This is very different from saying that their scale provides them some competitive advantage in providing care.
- bcoates 7y agoThis is a good guide. One thing I'd add: While you're monitoring for traffic/errors/latency throw in minimum success rate. Make a good estimate of how many successful operations monitored systems will do per minute on the slowest hour of the year and put in an alert if the throughput drops below that. You'd be surprised how many 0 errors/0 successes faults happen in a complex system, and a minimum throughput alarm will catch them.
- mkeedlinger 7y agoIn some systems it's also nice to have a "canary" acting as the user to call your API every minute or so. Then even during low / no traffic hours you can still catch errors/outages/etc.
- lrem 7y agoThat's a prober, not a canary.
- nitrogen 7y agoSometimes things can be described in more than one way. Canary, readiness probe, health check, heartbeat, etc. It makes communication more efficient and enjoyable when all parties are aware of this.
- lrem 7y agoYes, there are many words in the space. But there are also many classes of systems. It makes communication more efficient and enjoyable if all parties can think of the same thing when hearing a word.
- deleted 7y ago[deleted]
- VincentEvans 7y agoDid you do all these things by yourself? Really great content, but was really taken back by “I” used everywhere. Maybe it’s a new thing that I am not hip on that I ought to try - “I built and ran transaction processing software for Bloomberg! This is what I learned!” But perhaps you really did all that by yourself, in that case sorry that i doubted you, looks like it’s a lot.
- techie128 7y agoInteresting. Although it is on the lite side. For example, it doesn't talk about chaos testing, defining effective and comprehensive metrics (KPIs), alert noise or running services like databases in an active-active (hot-hot) mode.
- learnfromstory 7y agoDon't really agree that this list could have come about through discussions with engineers at Google, Facebook, etc. The more computers you have the less important it becomes to monitor junk like CPU and memory utilization of individual machines. Host-level CPU usage alerting can't possibly be a "must-have" if there are extremely large distributed systems operating without it. If you've designed software where the whole service can degrade based on the CPU consumption of a single machine, that right there is your problem and no amount of alerting can help you.
- madhadron 7y ago> If you've designed software where the whole service can degrade based on the CPU consumption of a single machine, that right there is your problem and no amount of alerting can help you. Unless it's your database.
- learnfromstory 7y agoIf you have "the database" then you're fucked anyway and probably your thing isn't on the scale that we are discussing.
- cameronbrown 7y agoIf you're using a sharded SQL database then a single machine going bad could still affect thousands of people.
- kevinsundar 7y agoI work at a FAANG and host level cpu is most definitely an alert we page on. Though a single host hitting 100% CPU isn't really a problem in and of itself (our SOP is just to replace the host), its an important sign to watch for other hosts becoming unhealthy. It might be overkill but hey theres mission critical stuff at hand. For example: if you have a fleet of hosts handling jobs with retries, a bad job could end up being passed host to host killing each host / locking up each one as it gets passed along. And that could happen in minutes while replacing and deploying and bootstrapping a new host takes longer. So by the time your automated system detects, removes, and spins up a new host everything is on fire.
- cpursley 7y agoCan anyone recommend MOOCs and/or university courses (open syllabus) covering Distributed Systems?
- jammygit 7y agoAlso curious! I've looked in the past and had a lot of trouble finding one. A book recommendation would also be very helpful!
- elamje 7y agoYeah, I just asked a similar question on HN and didn't get many responses, but one overwhelming book rec was "Designing Data Intensive Applications" Basically a high level guide through modern architectures, frameworks, and database designs. So far, my takeaway has been learning what tool would be useful for certain types of data engineering, not the details of how to write code with it. Edit - link: https://news.ycombinator.com/item?id=20417801 https://news.ycombinator.com/item?id=20417801
- dkersten 7y agoI second that book. Got a pre-release digital copy a couple of years back and its awesome. Physical copy is on my desk. I should make time to re-read it :)
- lazyant 7y agoUnfortunately doesn't seem to be good books out there or I don't know about them. The "Data intensive applications" book is highly praised but focused in databases. The other "Designing Distributed Systems:" O'Reilly book is solely about Kubernetes. The "Scalable Internet Architecture" is trash (sorry). The best source for distributed computing I found is Facebook's engineering blog.
- PhilippGille 7y agoSome links to related resources (repos, articles, books, courses) can be found here: - https://github.com/mxssl/sre-interview-prep-guide https://github.com/mxssl/sre-interview-prep-guide - https://github.com/theanalyst/awesome-distributed-systems https://github.com/theanalyst/awesome-distributed-systems One linked resource for example: - https://github.com/aphyr/distsys-class https://github.com/aphyr/distsys-class
- drdrey 7y agoI find it problematic that this recommends the Five Whys to get to "the root cause". Haven't we collectively moved past that?
- teraflop 7y agoWould you care to explain why you find that problematic?
- drdrey 7y agoSee this post by John Allspaw: https://www.oreilly.com/ideas/the-infinite-hows https://www.oreilly.com/ideas/the-infinite-hows > The Five Whys, as it’s commonly presented, will lead us to believe that not only is just one condition sufficient, but that condition is a canonical one, to the exclusion of all others. Five whys presents itself as a way to dig deep but promotes doing so linearly and getting to a singular thing you can fix, hiding a lot of potential learnings along the way. Thinking about contributing factors is a much more powerful framework.
- lrem 7y agoI find it useful: you'll find at least one problem deep down. The more surface problems will get worked on anyways. In any case, the point is to make your system more robust over time.
- james_s_tayler 7y agoThe notion of a "root cause" itself is flawed. It's causes acting in concert which cause problems and they all enable each-other.
- LaserToy 7y agoI would add - you should start your monitoring with business metrics. Monitoring low level things is good to have but putting whole emphasis on it is missing the whole point. You should be able to answer at any point of time whether users are having a problem, what problem, how many users, what are they doing to workaround? In other words, when person is in ER, doctors are looking for heartbeat, temperature, ... , not for some low level metric, like how many grams of oxygen is consumed by some specific cell.
- cperciva 7y agoYes, in the ER, doctors look at things like heart rate, respiration rate, and temperature. But they also draw blood for electrolytes, glucose, creatine kinase, etc. since those can detect underlying problems which the body is compensating for. A well designed distributed system is going to be able handle a certain failure rate in its components because requests will be retried automatically. If a component's failure rate increases from 0.01% to 0.1%, there will probably not be any user-visible impact... but if you can detect that increase, you might be able to correct the underlying issue before that component's failure rate increases to 1% or 10% or 100% -- at which point no amount of retrying will avoid problems.
- LaserToy 7y agoI;m not saying you should not do that, I'm saying your focus should not be there. When things go south your CEO rarely is interested in CPU utilization, or error rate. The question they usually ask: How bad is it? What is user impact? Is it all hand on deck or it is just a glitch. When it is cascading, which system to fix first? Root cause? Component failure rate just doesn't have enough context. And yes, distributed systems are hard, because it is inherently hard to reason about what will happen when something changes/fails. I'm not from Uber but I recently worked (was responsible for a huge chunk of infra) at a company that is bigger and has more products running and I saw some hilarious failures. And what is different, it was rarely a bad code push, and when it was, due to the nature of the business, sometimes it was really hard to roll back.
- GauntletWizard 7y ago
- NewsAware 7y agoNice article. The amount of in-house procedures/tooling developed by the backend seems impressive (maybe some not invented here syndrome but can't judge really). What I am astonished at though, is that the backend part of Über seems so professional while the Uber Android app feels like it's build by 2 junior outsourced devs. Have rarely used an app which felt so buggy and awkward. (aside from regular crashes, e.g. when registering for Jump-bikes in Germany inside the app, I had to restart the app to have the corresponding menue item appear).
- RhodesianHunter 7y agoReally? That's a bit surprising to me as I've always thought their app was slick, easy to use, and responsive. Out of curiosity what device are you on?
- jwilliams 7y agoGood read. There are a few things that I'd throw on top as important; - Active monitoring - Chaos testing - Cold start testing
- ggregoire 7y agoMost of those advices apply to small non-distributed systems too.
- toolslive 7y agoYou need to have a strategy for backward (and forward) compatibility for your components. If the environment is large enough, you don't exactly know which component is running what version of the code as they are constantly (holding back on) upgrading some part of your system. This includes extra (paramaters to) RPC calls, data type evolution, schema evolution. Without a decent strategy you'll be in over your head quickly. (Tip : a version number for your API as part of the API v0.0.1, ain't gonna be enough)