3 ms·
There is a lot to be said for "smoke test" metrics. Things you expect to have frequent false positives, but are sometimes early indicators of larger problems or
by jofer 2y ago
There is a lot to be said for "smoke test" metrics. Things you expect to have frequent false positives, but are sometimes early indicators of larger problems or indicators of where to look deeper if something else goes sideways. They're not things that should wake you up in the middle of the night, but they're a damn valuable tool to quickly figure out what's actually wrong when a "real" alert triggers.
Many of these lend themselves well to dashboards instead of alerts, but not everything is "dashboardable". Sometimes it's good to have a set of low-priority alerts that are treated differently than others.
E.g. "we're not receiving any data/requests". Sometimes that's just a lull in activity. Maybe a holiday. Sometimes it's because everything _else_ is broken and nothing is getting in (e.g. DNS issues).
With that said, I do think that classification should be made manually and not automatically.
- aflag 2y agoThat's why alerts can have different priority levels. So, less serious issues can be addressed during normal working hours (eg. disk is 80% full). Maybe LLMs will figure out the correct P level for something like disk usage, but it's unlikely to get it right for things that are particular to your application. Maybe use an LLM when you're creating the alert to auto fill the priority level. That can then be verified by someone. Don't silence an alert based on what an LLM thinks though.
- jofer 2y agoYeah, I completely agree. I just meant that alerts you don't immediately respond to or are "noisy" aren't necessarily things you want to delete. Having low priority "noisy" alerts is not a bad thing.