3 ms·
IDK about anomaly detection being on that list. When I worked at a large tech company, the in-house anomaly detection and root cause analysis capabilities worke
by Xcelerate 2y ago
IDK about anomaly detection being on that list. When I worked at a large tech company, the in-house anomaly detection and root cause analysis capabilities worked like magic compared to other places I've been. When done right, it can be extremely valuable.
- kitd 2y agoSounds like you hit the 1/10. Which is great and I agree, very rewarding.
- jldugger 2y agoHalf of the magic is knowing _when_ to apply it. If you just dump your prometheus timeseries data into an anomaly detection system and ask it to constantly scan for anomalies, you will always find them. As an SRE, I don't actually care about anomalies all that much. For initial alerting, I want phase shift detection. One customer sending a few bad API calls on one specific minute is uninteresting and pretty much inactionable. That same error rate over 10 minutes is more interesting and more likely to be a systems problem I can actually resolve. But the raw AD stream is just too damn noisy for all sorts of reasons. This is why our alerting tools have a `duration` field: any signal above threshold must remain so for multiple observation periods before summoning human inspection. And why health checks have grace periods and retries before killing services. Where anomaly detection works better, IMO, is post-alert analysis. At that point anomalies are welcomed as hypotheses, since the system features complex interactions between components. I've built a couple of dashboards using extremely simple math, like Laplace smoothing and time series correlation that help surface relevant information from the flood of metrics and logs collected. But critically, these tools generally don't use the time domain as their baseline. Usually, I'm comparing a cluster against another one in a different region, or one metric against another, rather than now versus twelve hours ago.