3 ms·
I'd think that receiving even 1/5th the rate limit in a 5 minute window would be disorienting enough to render alerting effectively useless. I'd question the c
by deeblering4 5y ago
I'd think that receiving even 1/5th the rate limit in a 5 minute window would be disorienting enough to render alerting effectively useless.
I'd question the configuration which fires that many alerts in that time frame, and suggest improving alert aggregations and dependencies to get the number down to one or a handful of meaningful alerts.
- curryst 5y agoThe overhead of maintaining those configurations all the time is usually too high to be worth it considering the benefit and likelihood of reaping it. Also, in my experience with those systems, they only make sense to use very sparingly. Your monitoring becomes extremely fragile when your aggregations and dependencies get complicated enough that "what will our alerting system do when X happens?" results in a flow chart with 18 steps. If you aren't careful, you can end up making your aggregations less useful than the raw alerts would be.
- rmetzler 5y agoIt would be great to have a dependency graph or labels in the alerts, so they are easily mapped to the things that can break and are important enough to be monitored. We just had a short outage where an editor removed the index page in the cms which is central to the site. It's stupid that this is possible but we just operate the cms while we build and operate everything around it for our customer. I think a large part of our alerts where triggered all at once but the one thing they had in common was that the alerts all pointed to the index page in the cms. E.g. the public www alert for index, the public api alert for index, the preview www alert for index, the preview api alert for index....
- CSDude 5y agoProblem is you get same alert deduplicated hundreds of time. And with those limits, you miss others.