5 ms·
It feels to me that using LLM to classify alerts as noisy is just adding risk instead of fixing the root cause of the problem. If an alert is known to be noisy
by aflag 2y ago
It feels to me that using LLM to classify alerts as noisy is just adding risk instead of fixing the root cause of the problem. If an alert is known to be noisy and have appeared on slack before (which is how the LLM would figure out it's a noisy alert), then just remove the alert? Otherwise, how will the LLM know it's noise? Either it will correctly annoy you or hallucinate a reason it figures that alert is just noise.
- aray07 2y agoyeah, thats the goal of adding the context and the report - to hopefully bring awareness to the team that this alert should be removed. My rationale for flagging the alert was to help prioritization for the on-call (lets say there are multiple alerts going off at the same time)
- ozim 2y agoThat’s a people problem and you cannot fix people problems with tech. If no one cares to do the good job of managing alerts putting AI in front of it will not change that.
- jrochkind1 2y agoAn AI could help bring to your attention alerts that need managing. I like it for this better than for someone in the moment of receiving an alert deciding whether or not to pay it attention.
- ozim 2y agoIf someone ignores alerts they will keep ignoring them but now you automated part of ignoring with "AI" and human at the end still will ignore alerts the same. Writing it out makes me laugh because that's like something from Douglas Adams stories. Automated Ignoring System along with Infinite Improbability Drive.
- jrochkind1 2y agoI was thinking of simply pointing out which kinds of alerts need to be tuned to be less noisy.
- digging 2y ago> you cannot fix people problems with tech For very specific values of "people problem", "fix", and "tech". In reality, a more true (and relevant) assertion is "appropriate tools can make virtually any problem more tractable." For example, it takes an annoyed engineer to notice that the same flaky alert keeps going off and is noise. Then it takes non-trivial skill on their part to communicate the need to disable that alert. They will meet non-trivial resistance, because disabling alerts is dangerous. However, if the tools they are using say "This is a noisy alert, it hasn't been useful for 6 months," disabling that alert becomes more of a best practice for the organization.
- jofer 2y agoThere is a lot to be said for "smoke test" metrics. Things you expect to have frequent false positives, but are sometimes early indicators of larger problems or indicators of where to look deeper if something else goes sideways. They're not things that should wake you up in the middle of the night, but they're a damn valuable tool to quickly figure out what's actually wrong when a "real" alert triggers. Many of these lend themselves well to dashboards instead of alerts, but not everything is "dashboardable". Sometimes it's good to have a set of low-priority alerts that are treated differently than others. E.g. "we're not receiving any data/requests". Sometimes that's just a lull in activity. Maybe a holiday. Sometimes it's because everything _else_ is broken and nothing is getting in (e.g. DNS issues). With that said, I do think that classification should be made manually and not automatically.
- aflag 2y agoThat's why alerts can have different priority levels. So, less serious issues can be addressed during normal working hours (eg. disk is 80% full). Maybe LLMs will figure out the correct P level for something like disk usage, but it's unlikely to get it right for things that are particular to your application. Maybe use an LLM when you're creating the alert to auto fill the priority level. That can then be verified by someone. Don't silence an alert based on what an LLM thinks though.
- jofer 2y agoYeah, I completely agree. I just meant that alerts you don't immediately respond to or are "noisy" aren't necessarily things you want to delete. Having low priority "noisy" alerts is not a bad thing.
- tasn 2y agoI like incident.io's take on LLMs with incident management: which is essentially assist, don't decide. [1] 1: https://5x9s.svix.com/p/evolution-of-incident-management https://5x9s.svix.com/p/evolution-of-incident-management