4 ms·
A major issue with on-call, and certainly one I've encountered multiple times, is the high likelihood of moral hazard - the people who are responsible for addre
by oofnik 4y ago
A major issue with on-call, and certainly one I've encountered multiple times, is the high likelihood of moral hazard - the people who are responsible for addressing incidents are not the same people who designed and maintained the system at fault. This results in the former team feeling powerless to put out fires which could have been prevented by more robust design, and the latter team having no incentive to improve reliability.
SRE gets this right, at least in theory, by requiring that all production systems be reviewed and approved, including observability and incident management procedures, prior to entering service. This ensures that there is some shared responsibility across teams for maintaining uptime.
https://sre.google/sre-book/being-on-call/ https://sre.google/sre-book/being-on-call/
- thesuperbigfrog 4y ago>> A major issue with on-call, and certainly one I've encountered multiple times, is the high likelihood of moral hazard - the people who are responsible for addressing incidents are not the same people who designed and maintained the system at fault. This results in the former team feeling powerless to put out fires which could have been prevented by more robust design, and the latter team having no incentive to improve reliability. The Amazon approach was (still is?) to have the team that develops and deploys the software to be responsible for the on-call rotation for that system. If you are developing software for such a team, it gives you a direct reason to make sure everything is designed and tested well before it is deployed to production--you (or a teammate) will be answering the early morning alert to fix it if it is not. A direct feedback loop like that is remarkably effective to prevent the moral hazard and ensures direct accountability when buggy software gets deployed.
- itsmemattchung 4y agoThis is still true today (at AWS). If you write the software, you are ALSO responsible for owning the operations: this means on call.
- jabroni_salad 4y agoThe most attractive thing to me in that SRE handbook is the error budget. I don't mind poking at your thingy but I'm only going to do it a few times before it becomes your problem instead.
- Aeolun 4y ago> SRE gets this right, at least in theory, by requiring that all production systems be reviewed and approved, including observability and incident management procedures, prior to entering service. Both doing, and being subjected to reviews suck balls. I’d much rather be on call 24/7 to fix the problems I caused myself.