4 ms·
The deep metrics, as you call them, get vanishingly irrelevant the larger the system gets, at least for the use case you mention. To put it unscientifically, s
by linza 7y ago
The deep metrics, as you call them, get vanishingly irrelevant the larger the system gets, at least for the use case you mention.
To put it unscientifically, something always goes wrong, but that does not mean SLA is violated. As a corrollary, you will be alerted constantly about something that might go wrong because of some signal that someone thought might lead to some outage.
At face value it seems proactive, but it's really not. Better spend time actively making the system more reliable (e.g. look at your or someone else's postmortems, or do premortem exercises).
- jorangreef 7y agoOn the contrary, low level metrics become more important the larger a system gets, since the quantity of low probability events (e.g. cosmic ray bit flips) increases. Without those "irrelevant" low level metrics for a large growing distributed system, you're not only flying blind but crashing increasingly more and more often.
- linza 7y agoYou are just chasing your own tail by looking at those metrics. Your system needs to be resilient to these random events and you must account for those, that's for sure. But you would not put your focus on those, it's still the higher level alerts that are more relevant to the business. Crashing more often is caught by an alert that looks at total capacity is s service. Doesn't matter if it's random bit flips or OOMing nodes at first. Longer term these metrics can be useful to increase efficiency again (should i first fix random bit flips or OOMing tasks). I would not base SLA relevant alerts on too low level alerts.
- jacquesm 7y agoThis is simply wrong. Quality starts in the basement, and works its way up. So if you monitor a system at the lowest level you will be able to build confidence about those layers and you will spot trends long before they will become apparent at higher levels because those higher levels will erase the fact that at lower levels things are already going wrong. Systems that only monitor the highest levels appear to function fantastically well right up to the moment they crash spectacularly. Then the forensics will show you that at lower levels there were plenty of warning signs telling you that the system was headed for the cliffs and a large number of those signs will be apparent while there was still time to do something about them. Reliable systems engineering is not something you can do just at the highest levels because of the build in resilience in intermediary levels. This is counter-intuitive but born out in countless examples of systems that look robust but aren't versus those systems that really are robust.
- antpls 7y agoAfter reading this thread, everyone convinced me there are good arguments for having different levels of metrics. I think we can better understand the problem if we introduce the word "priority". Priority is always the customer. The focus can change given the circumstances and the context of each issue. Sometimes the focus must be high-level, sometime it requires a low-level focus, but the priority is always the customer.