4 ms·
I would add - you should start your monitoring with business metrics. Monitoring low level things is good to have but putting whole emphasis on it is missing th
by LaserToy 7y ago
I would add - you should start your monitoring with business metrics. Monitoring low level things is good to have but putting whole emphasis on it is missing the whole point. You should be able to answer at any point of time whether users are having a problem, what problem, how many users, what are they doing to workaround?
In other words, when person is in ER, doctors are looking for heartbeat, temperature, ... , not for some low level metric, like how many grams of oxygen is consumed by some specific cell.
- cperciva 7y agoYes, in the ER, doctors look at things like heart rate, respiration rate, and temperature. But they also draw blood for electrolytes, glucose, creatine kinase, etc. since those can detect underlying problems which the body is compensating for. A well designed distributed system is going to be able handle a certain failure rate in its components because requests will be retried automatically. If a component's failure rate increases from 0.01% to 0.1%, there will probably not be any user-visible impact... but if you can detect that increase, you might be able to correct the underlying issue before that component's failure rate increases to 1% or 10% or 100% -- at which point no amount of retrying will avoid problems.
- LaserToy 7y agoI;m not saying you should not do that, I'm saying your focus should not be there. When things go south your CEO rarely is interested in CPU utilization, or error rate. The question they usually ask: How bad is it? What is user impact? Is it all hand on deck or it is just a glitch. When it is cascading, which system to fix first? Root cause? Component failure rate just doesn't have enough context. And yes, distributed systems are hard, because it is inherently hard to reason about what will happen when something changes/fails. I'm not from Uber but I recently worked (was responsible for a huge chunk of infra) at a company that is bigger and has more products running and I saw some hilarious failures. And what is different, it was rarely a bad code push, and when it was, due to the nature of the business, sometimes it was really hard to roll back.
- GauntletWizard 7y agoYou're missing the forest for the trees. The deeper metrics are important for diagnosis. The business metrics are what you care about. I'd phrase your advice differently as well: monitor on contact points. Monitor where rubber meets road. Monitor where two components interesect, be it between teams working on frontend and backend components, your application and it's database, or low level metrics - the system API. I've found many interesting problems and averted several potential crises because I saw that the metrics things like requests looked different between the client and server, the system metrics and my applications, etc. Contact points are what's important, and it just happens that every (Figuratively) application has a contact point with the OS.
- cperciva 7y agoBut that's my point: The deeper metrics aren't merely useful for diagnosis after the business metrics go sideways; they can be useful as leading indicators, to warn you before problems reach the point of affecting the business metrics which you care about.
- linza 7y agoThe deep metrics, as you call them, get vanishingly irrelevant the larger the system gets, at least for the use case you mention. To put it unscientifically, something always goes wrong, but that does not mean SLA is violated. As a corrollary, you will be alerted constantly about something that might go wrong because of some signal that someone thought might lead to some outage. At face value it seems proactive, but it's really not. Better spend time actively making the system more reliable (e.g. look at your or someone else's postmortems, or do premortem exercises).
- jorangreef 7y agoOn the contrary, low level metrics become more important the larger a system gets, since the quantity of low probability events (e.g. cosmic ray bit flips) increases. Without those "irrelevant" low level metrics for a large growing distributed system, you're not only flying blind but crashing increasingly more and more often.