3 ms·
I'm guessing the suggestion is: once the error budget is spent, no new feature work is accepted until causes of the errors are fixed
by bigbluedots 4y ago
I'm guessing the suggestion is: once the error budget is spent, no new feature work is accepted until causes of the errors are fixed
- kubanczyk 4y ago"Fixing the error" is at times subjective. A less nuanced approach is to freeze non-reliability-improving changes (i.e. merges) until the production meets SLO again. That is the canonical example policy given in Google's SRE Book.
- groby_b 4y agoSure, that's solid policy, but I'm not sure it really addresses the "tribal knowledge" issue per se. I mean, I'd like to think that fixing reliability issues increases knowledge, but decades on multithreaded codebases littered with "sleep(0); // just in case" point in a different direction. All I've found to work so far is a deliberate attempt to shift to more active knowledge sharing. I was hoping to learn a few new tricks from OP, but "reliability freezes" are not it, based on my experience.