5 ms·
> Alerts from each integration 300 5 minutes > Alerts from the whole team 500 5 minutes > API requests per API key 300 5 minutes Product looks great but thos
by CSDude 5y ago
> Alerts from each integration 300 5 minutes
> Alerts from the whole team 500 5 minutes
> API requests per API key 300 5 minutes
Product looks great but those API request limits are too low, because alerts rain when you are having incidents and rate limiting all of them is harmful. That's why other products have deduplication keys / aliases so you don't miss important ones.
https://grafana.com/docs/grafana-cloud/oncall/oncall-api-reference/ https://grafana.com/docs/grafana-cloud/oncall/oncall-api-ref...
- deeblering4 5y agoI'd think that receiving even 1/5th the rate limit in a 5 minute window would be disorienting enough to render alerting effectively useless. I'd question the configuration which fires that many alerts in that time frame, and suggest improving alert aggregations and dependencies to get the number down to one or a handful of meaningful alerts.
- curryst 5y agoThe overhead of maintaining those configurations all the time is usually too high to be worth it considering the benefit and likelihood of reaping it. Also, in my experience with those systems, they only make sense to use very sparingly. Your monitoring becomes extremely fragile when your aggregations and dependencies get complicated enough that "what will our alerting system do when X happens?" results in a flow chart with 18 steps. If you aren't careful, you can end up making your aggregations less useful than the raw alerts would be.
- rmetzler 5y agoIt would be great to have a dependency graph or labels in the alerts, so they are easily mapped to the things that can break and are important enough to be monitored. We just had a short outage where an editor removed the index page in the cms which is central to the site. It's stupid that this is possible but we just operate the cms while we build and operate everything around it for our customer. I think a large part of our alerts where triggered all at once but the one thing they had in common was that the alerts all pointed to the index page in the cms. E.g. the public www alert for index, the public api alert for index, the preview www alert for index, the preview api alert for index....
- CSDude 5y agoProblem is you get same alert deduplicated hundreds of time. And with those limits, you miss others.
- named-user 5y agoHow else do you think they are gonna make money?
- CameronNemo 5y agoThat's why other products have deduplication keys / aliases so you don't miss important ones. Care to link to the docs? I'm interested.
- CSDude 5y agohttps://support.atlassian.com/opsgenie/docs/what-is-alert-de-duplication/ https://support.atlassian.com/opsgenie/docs/what-is-alert-de... https://support.pagerduty.com/docs/event-management https://support.pagerduty.com/docs/event-management
- CameronNemo 5y agoThanks for the links. From the article: With Grafana OnCall’s automatic grouping of alerts within Slack, you can avoid alert storms and reduce the noise your teams are exposed to during an incident. Seems like the same feature described using different terminology.
- EwanToo 5y agoThe output alerts feature looks largely the same, but the input API limits are the part in question. What happens if you get 1000 API calls about "Alert 1" and 1 API call about "Alert 2". You want both on call's to trigger once, but will alert 2 get though?
- dharmab 5y agoI was once in a job where I was solo on call for tens of thousands of cores globally and at worst we had like 2000 alerts in a week. These limits seem quite high to me.