4 ms·
Blameless culture should be a standard in the engineering industry
- jkic47 3y agoMany years ago, I ran a site that used several 100 K kg of a toxic, explosive and flammable gas. They were just about to hit their permitted levels of emissions and didn't know why. I sat with the team and said to them that everything in the past was done with the knowledge available at the time - water under the bridge. We set ourselves small goals and worked our way through the system from beginning to end, doing little experiments to prove/disprove our assumptions. Over the timescale of months, the attitudes changed, and my engineers began to say things like "we were wrong when we did this/that". We had just one rule - if the data shows that a previously made assumption was wrong, accept it and fix it. We went from ~1000 kg emissions down to the ppb level over time. One of the best teams I had worked with, and a great memory.
- rhelz 3y agoI question whether real human beings, in a real work environment, can actually rise to the angelic level of not blaming anybody. Some guy causes you to spend 10 all nighters in a row, because every day for 10 days he moved bad code to production? Are you seriously expecting me not to start assigning blame? If I were that guy, should I really expect that come performance review time, that would be completely ignored? I also question whether it's even a good idea, because if you throw out the blame, you throw out the accountability as well. I'm a strong believer in giving senior devs a very large degree of freedom in how they code up solutions. But the more discretion somebody has, the higher should be the standard to which they are held. Nevertheless, everybody makes mistakes, so what the heck are you supposed to do? What I'd suggest instead is to create an environment where the consequences for making a mistake are as small as possible. The architecture of the code should allow for incremental improvements. There should be extensive unit tests. The process of going from development to production should catch as many mistakes as possible. And even when things are deployed to production, negative consequences should be minimized. If you have, say, 10 processes for accepting orders, only upgrade one of them at a time. If the new code crashes in prod, its ok because the other 9 redundant processes can step up. If you have 10 servers balancing the workload, only upgrade one server at a time. Etc Etc. Everybody makes mistakes. What's more, sometimes you need to take risks on new libraries or new techniques. No risk no reward. But if you design your architecture and build&release with mistake minimization in mind at all levels, you can make those mistakes and take those risks with minimal bad consequences.