10 ms·
I don't know about others, but I can't help but smile when I read the detailed series of events in aviation postmortems. To be able to zero in on what turned ou
by modernpacifist 3y ago
I don't know about others, but I can't help but smile when I read the detailed series of events in aviation postmortems. To be able to zero in on what turned out to be a single faulty part and then trace the entire provenance and environment that led to that defective part entering service speaks to the robustness of the industry. I say that sincerely since mistakes are going to happen and in my view robustness has less to do with the number of mistakes but how one responds to them.
Being an SRE at a FAANG and generally spending a lot of my life dealing with reliability, I am consistently in awe of the aviation industry. I can only hope (and do my small contribution) that the software/tech industry can one day be an equal in this regard.
And finally, the biggest of kudos to the Kyra Dempsey the writer. What an approachable article despite being (necessarily) heavy on the engineering content.
- sylens 3y agoI think many of us are so used to working with software, with its constant need for adaptation and modification in order to meet an ever growing list of integration requirements, that we forget the benefits of working with a finalized spec with known constants like melting points, air pressure, and gravity.
- abid786 3y agoCompletely agree - I think it can go one of two ways. Software is more malleable than airplanes are and that also comes with downsides (like how much time and effort it takes to bring a new plane to the market)
- RajT88 3y agoI was just thinking of this metaphor today. Try drawing the software monstrosity you work on / with as an airplane. 100 wings sticking out all different directions, covered with instruments and fins, totally asymmetrical and 5 miles long. Propellers, jets, balloons, helicopter blades. Yep, it flies. When it crashes, just take off again.
- twothamendment 3y agoSo software is my son's Bad Piggies flying monstrosity! You only left out the crates of TNT.
- otherme123 3y agoThe article talks about a piece of software that partially failed, when they needed to calculate the braking distance for the overweight aircraft.
- WalterBright 3y agoAirliners face constantly changing specifications. No two airliners are built the same.
- spenczar5 3y agoDo you mean no two individual planes? Like two 767s made a month apart, do you mean they literally would have different requirements?
- MBCook 3y agoI think they meant a 737-400 is different from a 737-500 is different from a 787 and a AirBus 320 and a MD-80 and… Every single model is somewhat bespoke. There’s common components but each ends up having its own special problems in a way I assume different car models in a common platform (or two small SUVs from competing manufacturers) just don’t.
- WalterBright 3y agoYes. There are constant changes to the design to improve reliability, performance, and fix problems, and the airlines change their requirements constantly.
- networkchad 3y ago[dead]
- ponector 3y agoI think they means that airplanes are made in different versions, catered to particular airline. Also planes are constantly updated. Two 767 made few months apart will have initial difference, like two different versions of java 8 SDK.
- numpad0 3y agoNeat little detail of the world Wikipedia once told me: the 00 suffix of classic Boeing planes, dropped in 2016, was substituted with Boeing assigned customer code on registration documents. e.g. a PAN AM 773-300 would have been 777-321, an Air Berlin Jetfoil would have been 929-16J, and so on. 1: https://en.wikipedia.org/wiki/List_of_Boeing_customer_codes https://en.wikipedia.org/wiki/List_of_Boeing_customer_codes
- nextos 3y agoAviation is great because the industry learns so much after incidents and accidents. There is a culture of trying to improve, rather than merely seeking culprits. However, I have been told by an insider that supply chain integrity is an underappreciated issue. Someone has been caught selling fake plane parts through an elaborate scheme, and there are other suspicious suppliers, which is a bit unsettling: "Safran confirmed the fraudulent documentation, launching an investigation that found thousands of parts across at least 126 CFM56 engines were sold without a legitimate airworthiness certificate." https://www.businessinsider.com/scammer-fooled-us-airlines-by-selling-fake-engine-parts-filings-2023-10 https://www.businessinsider.com/scammer-fooled-us-airlines-b...
- EdwardDiego 3y agoAdmiral Cloudberg has covered a case where counterfeit or EOL-but-with-new-paperworks components were involved in a crash. https://admiralcloudberg.medium.com/riven-by-deceit-the-crash-of-partnair-flight-394-f8a752f663f8 https://admiralcloudberg.medium.com/riven-by-deceit-the-cras...
- inglor_cz 3y agoI suspect this is precisely what is happening in Russian civil aviation now. No legit parts supplied, so there will be a lot of fake/problematic parts imported through black channels.
- crabmusket 3y ago> To be able to zero in on what turned out to be a single faulty part and then trace the entire provenance and environment that led to that defective part entering service speaks to the robustness of the industry. And to be able to reconstruct the chain of events after the components in question have exploded and been scattered throughout south-east Asia is incredible.
- Gare 3y agoMy impressiom was that the defective part was still inside the engine when it landed.
- EdwardDiego 3y agoProbably a reference to other incidents. Shout out to the NTSB for fighting off alligators while investigating this crash... https://en.wikipedia.org/wiki/ValuJet_Flight_592 https://en.wikipedia.org/wiki/ValuJet_Flight_592
- d1sxeyes 3y agoMakes it even more impressive: the parts that were actually implicated in the explosion itself (and scattered from the aircraft) were not defective, so the investigation had to go through parts which did not seem to have exploded in order to track down the defect. Or at least, I assume the turbine parts weren’t defective, although given what seems to be quite a happy-go-lucky approach to manufacturing defects in Hucknall, maybe my assumption is not made on solid grounds…
- Horffupolde 3y agoIf 200 people died after a db instance crashed, software would be equal in that regard.
- girvo 3y agoTo prove this, software that deals with medical stuff is somewhat more like aviation.
- cwalv 3y agoAlso, aviation and software aren't orthogonal. E.g., the article mentioned that part of the reason the pilot was able to sustain a very narrow velocity window between stall and overrunning the runway was because of the A380's fly by wire system.
- conradev 3y agoYep. Insulin pumps can kill their owner and the software updates need to be FDA approved: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4773959/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4773959/
- mlrtime 3y agoLikewise, in "aviation" when the entertainment system completely fails in a 4 hour flight, there is most like no post mortem at all. They turn it off/on again just like most of us.
- baby_souffle 3y agoThis is true in a lot of industries. Unless there’s 7+ figure costs or significant human losses, there’s usually not an exhaustive investigation to conclusively point to the exact cause and chain of events.
- mewpmewp2 3y agoSome people who think this is ideal for any sort of software tech sound they would also want a 3 hour post mortem with whoever designed the rooms, after slightly stubbing a toe.
- colechristensen 3y agoAerospace things have to be like this or they just wouldn’t work at all. There are just too many points of failure and redundancy is capped by physics. When there’s a million things which if they went wrong could cause catastrophic failure, you have to be really good at learning how to not make mistakes.
- WalterBright 3y ago> you have to be really good at learning how to not make mistakes. Not exactly. The idea is not not making mistakes, it's whatcha gonna do about X when (not if) it fails.
- WalterBright 3y agoAs a former Boeing engineer, other industries can learn a great deal from how airplanes are designed. The Fukushima and Deepwater Horizon disasters were both "zipper" failures that showed little thought was given to "when X fails, then what?" Note I wrote when X fails, not if X fails. It's a different way of thinking.
- cedivad 3y ago> When my AoA sensor fails, then what? crickets, let's just randomise which sensor we use during boot, that ought to do it!
- uselpa 3y agoEpic fail indeed, costing many lives.
- rytis 3y ago"AoA sensor" - Angle of Attack sensor. And the reference is presumably to 737 MAX accident. https://www.afacwa.org/the_inside_story_of_mcas_seattle_times https://www.afacwa.org/the_inside_story_of_mcas_seattle_time...
- asystole 3y ago> Airlines really want to be able to use pilots' existing type-rating on this hulking zombie of a 60s-era airframe with modern engines but it behaves differently under certain conditions, what do we do? let's just build a system that pushes the nose down under those conditions, have it accept potentially unreliable AoA data, and not tell pilots about it!
- f1shy 3y agoAs an engineer I think a lot about tradeoffs of cost vs other criteria. There is little I can learn from nuclear or aviation industry, as the cost structure ist so completely different. I’m very happy that the costs of safety in aviation are very good accepted, but I understand that few people are willing to pay similar costs for other things like, say, cars.
- mzi 3y agoIt took hundreds of subject experts from ten organizations in seven countries almost three years to reach that conclusion. Here at HN we want a post mortem for a cloud failure in a matter of hours.
- modernpacifist 3y ago> Here at HN we want a post mortem for a cloud failure in a matter of hours. I'll go one further - I've yet to finish writing a postmortem on one incident before the next one happens. I also have my doubts that folks wanting a PM in O(hours) actually care about its contents/findings/remediations - its just a tick box in the process of day-to-day ops.
- bitcharmer 3y agoApples to oranges
- thaumasiotes 3y agoSomething similar that struck me was that, in early February, Russia invaded Ukraine. And then, I saw an endless stream of aggrieved comments from people who were personally outraged that the outcome, whatever it might be, hadn't been finalized yet at the late, late date of... late February.
- mlrtime 3y agoI work at mid tier FAANG, our SLA for post mortems have SLA in the 7-14 day period. Nobody seriously wants a full PM in hours. They may want a mitigation or RCA in hours, but even AWS gives us NDA restricted PMs in > 24 hours.
- switch007 3y ago> I can only hope that the software/tech industry can one day be an equal in this regard I’d love to be an engineer with unlimited time budget to worry about “when, not if, X happens” (to quote a sibling comment). But people don’t tend to die when we mess up, so we don’t get that budget.
- solids 3y agoI agree, and also I enjoy the attitude. While in my profession the postmortems goal is finding who to blame, here the attitude is towards preventing it to happen again, no matter what. Or at least that’s how I feel.
- mewpmewp2 3y agoYour profession? Or you mean your company? Unless it's a very specific profession I would not know, it would usually imply that the company is dysfunctional.
- jstanley 3y ago> robustness has less to do with the number of mistakes but how one responds to them It must have something to do with the number of mistakes, otherwise it's all a waste of time! It's all well and good responding to mistakes as thoroughly as possible, but if it's not reducing the number of mistakes, what's it all for?
- krisoft 3y ago> It must have something to do with the number of mistakes, otherwise it's all a waste of time! Not really. Imagine two systems with the same amount of mistakes. (Here the mistakes can be either bugs, or operator mistakes.) One is designed such that every mistake brings the whole system down for a day with millions of dollars of lost revenue each time. The other is designed such that when a mistake happens it is caught early, and when it is not caught it only impacts some limited parts of the system and recovering from the mistake is fast and reliable. They both have the same amount of mistakes, yet one of these two systems is wastly more reliable. > if it's not reducing the number of mistakes, what's it all for For reducing their impact.
- bambax 3y agoThe Checklist Manifesto (2009) is a great short book that shows how using simple checklists would help immensely in many different industries, esp. in medical (the author is a surgeon). Checklists of course are not the same as detailed post-mortems but they belong to the same way of thinking. And they would cost pretty much nothing to implement. Also CRM: it's very important to have a culture where underlings feel they can speak up when something doesn't look right -- or when a checklist item is overlooked, for that matter.
- sgarland 3y agoYes, but they do have one critical failure mode: that the checklist failed to account for something (or that an expected reaction to a step being performed didn’t occur). I was a submarine nuclear reactor operator, and one of my Commanding Officers once ordered that we stop using checklists during routine operations for precisely this reason. Instead, we had to fully read and parse the source documentation for every step. Before, while we of course had them open, they served as more of a backstop. His argument – which I to some extent agree with – was that by reading the source documentation every time, we would better engage our critical thinking and assess plant conditions, rather than skimming a simplified version. To be clear, the checklists had been generated and approved by our Engineering Officer, but they were still simplifications.
- andrewaylett 3y agoIf the alternative to the check list is reading the full documentation, that's one thing. But in my experience -- as a Software Engineer, and random dude on the Internet -- the alternative is usually no check list or documentation.
- sgarland 3y agoFor sure – short of large and well-supported projects like Django et al., docs are notoriously incomplete if present at all. Even then, you have to get people to read them, which is somehow a monumental task. Docs? Nah, lemme read this Medium blog instead.
- blauditore 3y agoThis kind of makes sense, but it is only possible because of public pressure/interest. Many people are irrationally emotional about flying (fear, excitement etc.), that's why articles and documentaries like this post are so popular. On a side note, that's also why there's all the nomsense security theater at airports.
- Simon_ORourke 3y agoA colleague of mine came from a major aviation design company before joining tech and said they were in a state of culture shock at how critical systems were designed and monitored. Even if there are no hard real time requirements for a billing system, this guy was surprised at just how lax tech design patterns tended to be.
- mewpmewp2 3y ago> Being an SRE at a FAANG and generally spending a lot of my life dealing with reliability, I am consistently in awe of the aviation industry. I can only hope (and do my small contribution) that the software/tech industry can one day be an equal in this regard. There's a slight difference in terms of what kind of damage an airplane malfunctioning causes compared to a button on an e-commerce shop rendering improperly for one of the browsers. My point is that the level of investment in reliability and process should be proportional to the potential damage of any incidents.
- akarve 3y agoHard agree. Civil & mechanical engineering have a culture and history of blameless analysis of failure. Software engineering could learn from them. See the excellent To Engineer is Human in just this topic of analyzed failures in civil engineering.
- bomewish 3y agoRichard Hipp talks a lot about how SQLite adopted testing procedures directly from aviation.