4 ms·
The fact that an individual can bring production down is 1000% a problem in the process; outages should not attributed to single commits but to a series of prob
by hackandtrip 5y ago
The fact that an individual can bring production down is 1000% a problem in the process; outages should not attributed to single commits but to a series of problems in the process.
Just saying this since the number of times someone being production down should not be cause of shame or firing
- kempbellt 5y agoSomeone has to have the keys to the kingdom. Ideally, multiple people do - bus factor. If no one can bring production down this is also a problem. If only one person can, better hope they don't get hit by a bus, or take a vacation, or rage quit.
- hackandtrip 5y ago100% agree on this and that's my point. Was that code merged without any review? If so, that is a recipe for disaster. Why no unit tests, integration tests, QA? What was the speed at which the outage was resolved? If a sane deployment strategy is in place, I'd hope that no big damage was actually done; if the error budget was burning fast, there should be some manual/automatic rollback. tl;dr they may have been a key factor in the outage, but it seems like it is a symptom of a deeper hole in the process.
- Sparkyte 5y agoIt is complicated. They released code out of cycle which isn't a process issue. It's a person issue. You don't do that stuff without planning a MCM. You also don't become unreachable after a deploy.