2 ms·
> Teams never seem to understand how to alert on stuff. Ive been paged for things going off, that might indicate a problem, then you get stuck sticking around b
by chronid 4y ago
> Teams never seem to understand how to alert on stuff. Ive been paged for things going off, that might indicate a problem, then you get stuck sticking around because someone else wants to just wait and see what happens. "We should just be cautious" Its impossible to push back on these things, your just going against someones gut feeling, like maybe one day we will want to know, and everyone needs to protect them selves.
From an ops person: if an alert does not have:
- clear, provable impact on customers (internal/external)
- clear documentation (e.g. runbooks) on how to solve it
It should not be an alert. I took this path (successfully) when trying to remove spurious alerts that existed only for the ego of someone (most absurd example, something that started complaining when p99 for some endpoints went >500ms and happened every day when we downscaled the ASGs because business hours were over. No clear path to resolution, and impact was a couple pages opened a bit slower sure - but the number of customers using those pages after hours was <1%!
It sucks, definitely, but the best way to go around those alerts is to prove they're pointless or a waste of time or can be automated around and should automated around (and I've seen so many servlets leaking memory triggering OS alerts for OS teams or spawning infinite threads and never cleaning up after themselves...).
If the company does not want to do it, and pushes back, I would recommend starting to look for another company. It's sad, but it is what it is. 99.99% of software does not need a follow the sun rotation (or people damned to night shifts), just a bit of thought about what happens when things fail.