6 ms·
I think there's a lot to learn from the aviation industry. I did a talk at my companies internal conference on this (turned into words at https://medium.com/ing
by jefffoster 7y ago
I think there's a lot to learn from the aviation industry. I did a talk at my companies internal conference on this (turned into words at https://medium.com/ingeniouslysimple/why-dont-planes-crash-14a0579a5e2d https://medium.com/ingeniouslysimple/why-dont-planes-crash-1...).
For me it's the mindset that differs. Too often as software engineers we find a bug and just fix it. Aviation goes a step deeper and finds the environment that created the bug and stops that.
Unfortunately, the recent 737 MAX incidents seem to have changed this. From what I understand the reaction to the problems sounds more like what I'd expect a software business to do, rather than the airline industry!
- 0x445442 7y agoThe blog post was good but it was just another variation highlighting the age old conundrum... fast, good and cheap, pick two. Those that value quality are going to be swimming up stream in most organization that develop software because the bean counters always go straight to fast and cheap.
- maxxxxx 7y agoI work in regulated industry and you really don’t want the level of scrutiny regulated processes have in other industries. Innovation would slow down to a crawl or pretty much stop. In a lot of industries you can make trade offs quality vs speed or innovation and be better off by not having perfect quality.
- retiredcoder 7y agoI strongly agree. However to my experience companies decide on quality vs speed based on other factors rather than business needs. For example, in my previous job, CTO called “ overengineering” whenever something went against his will, while “innovation/we are not a Corp” was his tail wind. So much office politics for such a small company :/
- mcguire 7y ago...until you start injuring people or costing significant numbers of zeros.
- ellius 7y agoAfter fixing a recent bug, I asked my client company what if any postmortem process they had. I informally noted about 8 factors that had driven the resolution time to ~8 hours from what probably could have been 1 or 2. Some of them were things we had no control over, but a good 4-5 were things in the application team's immediate control or within its orbit. These are issues that will definitely recur in troubleshooting future bugs, and doing a proper postmortem could easily save 250+ man hours over the course of a year. What's more, fixing some of these issues would also aid in application development. So you're looking at immediate cost savings and improved development speed just by doing a napkin postmortem on a simple bug. I can't imagine how much more efficient an organization with an ingrained and professional postmortem culture would be.
- qznc 7y agoFor anybody into podcasts, I can recommend "Causality" https://engineered.network/causality/ https://engineered.network/causality/ John Chidgey digs into well known catastrophes, analyses what went wrong, and what was fixed afterwards. Not software related but promotes a safety mindset very well.
- warrenm 7y agoI love Chidgey's podcasts
- JohnChidgey 7y ago:heart:
- jasode 7y ago>as software engineers we find a bug and just fix it. [...] Unfortunately, the recent 737 MAX incidents seem to have changed this. I think there's some nuance about MCAS that's lost in all the media reports. As far as I understand, the MCAS software didn't have a "bug" in the sense we programmers typically think of. (E.g. Mars Climate Orbiter's software programmed with incorrect units-of-measure.[0]) Instead, the MCAS system was poorly designed because of financial pressure to maintain the fiction of a single 737 type rating. In other words, the MCAS software actually did what Boeing managers specified it to do: 1) Did the software only read a _1_ AOA sensor with a single-point-of-failure instead of reading _2_ sensors? Yes, because that was what Boeing managers wanted the software to do. It was purposefully designed that way. If the software was changed to reconcile 2 sensors, it would then lead to a new "AOA DISAGREE" indicator[1] which would then raise doubts to the FAA that Boeing could just give pilots a simple iPad training orientation instead of expensive flight-sim training. Essentially, Boeing managers were trying to "hack" the FAA criteria for "single type rating". 2) Did software make adjustments of an aggressive and unsafe 2.5 degrees instead of a more gentle and recoverable 0.6 degrees? Yes, because Boeing designed it that way. Somebody at Boeing specified the software design to be "1 sensor and 2.5 degrees" and apparently, that's what the programmers wrote. I know we can play with semantics of "bug" vs "design" because they overlap but to me this seems to be a clear case of faulty "design". The distinction between design vs bug is important to let us fix the root cause. The 737 MAX MCAS software issue isn't like the Mars Climate Orbiter or Therac-25 software bugs. The lessons from MCO and Therac-25 can't be applied to Boeing's MCAS because that unwanted behavior happens in a layer above the programming: - MCO & Therac: design specifications are correct; software programming was incorrect - Boeing 737MAX MCAS: design specifications incorrect; software programming was "correct" -- insofar as it matched the (flawed) design specifications [0] https://en.wikipedia.org/wiki/Mars_Climate_Orbiter#Cause_of_failure https://en.wikipedia.org/wiki/Mars_Climate_Orbiter#Cause_of_... [1] yellow "AOA Disagree" text at the bottom of display: https://www.ainonline.com/sites/default/files/styles/ain30_fullwidth_large_2x/public/uploads/2019/03/aoa_vaneindicator_aoa_disagree_lg2.jpg?itok=vNHrwopw https://www.ainonline.com/sites/default/files/styles/ain30_f...
- Kurtz79 7y ago"If the software was changed to reconcile 2 sensors, it would then lead to a new "AOA Disagree" indicator which would raise doubts to the FAA that Boeing could just give pilots a simple iPad training orientation instead of expensive flight-sim training." I always liked this quote from the "Mythical Man-Month": “Never go to sea with two chronometers, take one or three”. https://blog.ipspace.net/2017/01/never-take-two-chronometers-to-sea.html https://blog.ipspace.net/2017/01/never-take-two-chronometers...
- ken 7y agoThere are a handful of highly respected books that everyone knows software engineers should read, like "The Mythical Man-Month" and "Peopleware". Yet whenever I read one of these, I found I learned very little. Everything in them was obvious -- to those who are 'in the trenches'. What we need is a way to get managers to read these, and take them to heart. Even when my manager had a copy of the book sitting on their desk, they rarely had read it, and they absolutely never followed its advice. (When pushed, they might say "That was a groundbreaking book, for its day, but the industry moved on." Now we've got open floor plans, and AGILE SCRUM, and free snacks ... and also no evidence these are an improvement to the software development process, but never mind.) This aviation mindset you refer to is the same way. I can't tell you how many times this happened to me: - User clicks a button, and it doesn't do what it says it should. - A bug is filed, and assigned to me. - I investigate, and find the problem. I start preparing a fix. - Manager comes by to pester me. "Why isn't this button fixed? Shouldn't that have been a quick fix?" We played Planning Poker last week and everybody else who isn't working on it agreed it should only be a 1! - See, we're computing this value incorrectly, and I grepped the codebase and it turns out we're also doing it wrong in 7 other places, which causes... - "The customer wants this one button fixed. Don't worry about the others. Don't worry about testing, or cleaning up, or documenting why the mistake was made or how it should have been done. Those aren't on this milestone. Just fix this one button and move on. We need you working on the new features we promised our customers this month..." Modern software development is a circus of improperly aligned incentives.
- asn1parse 7y agoWhen they did the open space with free snacks I told them I was never coming into the office again. I pointed out that I got sick roughly 3 times a year from just being forced to spend time in that hazardous office. And I then I sealed it by telling them that they will get 2 extra hours of work out of me each day plus all the CO2 I wouldnt be pumping into the atmosphere each day when I drive there on the perilous freeways with all the other pissed off people on the road. And they bought it. This was 2012. Ffw to 2019, I still work at home for the same company. I never get sick, I dont have to deal with any of the toxic culture stuff at work, and I buy and consume my own snacks, the ones I want to eat. I agree, it's a circus and it's a lot easier to manage from a distance.
- qznc 7y agoI work in automotive. In Europe there is the ASPICE standard which is actually a reasonable guideline for (commercial) software development (unit tests, code reviews, etc). Customers require you to follow it. Top management requires you to follow it. Projects still ignore it. Writing unit tests at the end of a project misses most of the point, for example.
- qznc 7y agoThis was a nice prompt for me to finish my ASPICE article: http://beza1e1.tuxen.de/aspice.html http://beza1e1.tuxen.de/aspice.html
- jayd16 7y agoRetrospectives are a common part of agile. Only slightly less common is skipping retrospectives.