3 ms·
> alerts caused by issues in downstream systems that you could do nothing about How does Google deal with issues caused by downstream systems causing alerts?
by newman123 4y ago
> alerts caused by issues in downstream systems that you could do nothing about
How does Google deal with issues caused by downstream systems causing alerts?
- okdood64 4y agoGenerally 1) Mitigate if possible from your service's end, while simultaneously 2) paging the dependent service's team to mitigate/resolve; and if this happens too frequently or if it was a particularly bad incident you can push the other team to 3) create a postmortem with follow up AIs if they already didn't do so.
- cletus 4y agoThe deeper you go into the stack, the more reliable things tend to get, the more mature those systems tend to be and the more likely they are supported by SRE who take things very seriously. So if you're having an issue with Spanner, first it's likely not a bug in spanner. If it's an outage, somebody has probably already been paged. But if not, paging someone responsible will be answered quickly and treated seriously. You could've unexpectedly gone over quota on something. More often than not you can alleviate that with temporary quota while you resolve your issue (by reducing your usage, getting more permanent quota or both). A big part of this is that it's a cultural thing.