8 ms·
The failure was in a technical organization architecting a solution that allowed a simple human error in data entry to take down the entire system. This is clas
by qbasic_forever 4y ago
The failure was in a technical organization architecting a solution that allowed a simple human error in data entry to take down the entire system. This is classic system design and process failure. Bad input should have been caught before it got pushed to prod, before it got pushed to backup databases, etc. There was clearly no testing, no change process, no backup verification, etc.
People in senior management positions are at fault for creating a broken org and product. At a private company a failure like this would see the CTO immediately fired and even the CEO having to beg the board for forgiveness and keeping their job. The fault is not with the two contractors, it's with the bad engineering that enables a mistake to destroy the product.
- sokoloff 4y agoI agree that the ultimate root fault isn’t the recent data entry. > At a private company a failure like this would see the CTO immediately fired and even the CEO having to beg the board for forgiveness and keeping their job. That part doesn’t match my experience at all.
- toomuchtodo 4y agoRight. Engineer closest to the fault gets scapegoated and fired, new hire gets hired in without any idea what they’re coming into, and the show goes on. Good interview question to ask: “am I replacing a recent departure? If so, why?” You either get the truth or deception signal once hired. Both are helpful to know, and asking the question is free.
- marcosdumay 4y ago> Engineer closest to the fault gets scapegoated and fired That happens. But not every time, and even "usually" requires some statistics collecting to be sustained. Very often people just say "yes, we learned an expensive lesson", and change something that may or may not help the next time it happens.
- toomuchtodo 4y agoReally just depends on the org, and you’re not going to know until you’re on the inside.
- marcosdumay 4y agoYep, I completely agree with that. It can change from team to team too. Also, it can suddenly change at any time without a warning.
- hermitdev 4y agoI think it also matters on what the response to the incident was, in particular the answer to the questions of: who was responsible and how can we prevent this from happening again? If the answer to the first question is a bunch of finger pointing or the answer to the second is along the lines of "nothing", heads will likely roll. My first boss as a FTE at a trading shop made a change to a production database outside of process before a weekend. Come Monday morning, production trading systems couldn't come up for market open. We were also an options market maker. He refused to accept responsibility for his actions and denied there was anything done to prevent it. He was quickly behind closed doors in his boss's office. I had to undo the changes he made the week before (hooray for audit tables) as soon as I could. We were potentially liable for a substantial penalty from the exchange for being out of the market, and every minute mattered. Don't know if we ever did get penalized, but he was fired and out the door by 10AM that day. I've certainly seen and made lots of mistakes over the decades since, but that's the only one for which I've seen someone canned. Response matters.
- drstewart 4y ago>Right. Engineer closest to the fault gets scapegoated and fired, new hire gets hired in without any idea what they’re coming into, and the show goes on. That part doesn’t match my experience at all.
- pleb_nz 4y agoYou can just about tell what parts of the world a poster comes from when they start throwing the word 'fire' around willy nilly. Good question though..
- KennyBlanken 4y agoYou can tell what part of the world a poster comes from when they're smugly pretending that companies still don't have plenty of levers to pull to push you out the door, doing so by making you miserable. Oh, and the rampant racism and classism, thanks to the standard practice of including a photo on one's resume...
- scruple 4y agoSame here. In the places I've experienced with poor management and engineering cultures, what I've seen is management dictates all sorts of horrible, asinine things (like story points being treated as deadlines, etc.) that leads directly to problems like this because engineering is not given the space it needs to create robust solutions in the first place. The prevailing winds are, "Just get it done." The outcome is constant fire fighting and triage because of simple problems like bad data finding it's way into the system.
- drbeast 4y ago[dead]
- bennyelv 4y agoIt might not be the CTO's fault at all - maybe they inherited it, identified the issues, and are getting them solved. Probably not in this case, but you couldn't assume for every...
- chiefalchemist 4y agoFunny. I'm reading The Phoenix Project and just finished the about the CEO chewing out various employees. I keep thinking, responsiblity flows upstream, or should. Ultimately it is the CEO's job to put his people and the project in a position to suceed. Maybe the CEO in the book gets his due? Blaming a contractor? To save face? Isn't saving face. It's an embarrassment.
- noughtme 4y agoStill looks interesting, but I was hoping The Phoenix Project was a post-mortem of the Phoenix Pay System (scandal) https://en.m.wikipedia.org/wiki/Phoenix_pay_system https://en.m.wikipedia.org/wiki/Phoenix_pay_system
- chiefalchemist 4y agoIt may be "inspired by" as in the book TPP is among other things about the stores POS. TBH IDK the book was mentioned on HN and I figured I'd check it out.
- fineIllregister 4y ago> That part doesn’t match my experience at all. Exactly! There wasn't any real fallout for Southwest, who had a shutdown less than a month ago, for similar outdated technology reasons.
- Someone1234 4y agoMaybe even more pertinent is that they had a similar one in 2016 (smaller scale, but similar cause/effect) and then didn't fix anything, so it could happen again recently. There's running lean and then there is running recklessly.
- CamperBob2 4y ago"Move slowly and break things!"
- StreamBright 4y agoYeah you are right. Also, there is another side of this coin. I used to work for a FAANG and here is an outage story for you. We had every angle covered of preventing user errors and stopping a single data entry to take down systems. We had reviews of changes, multiple approval required etc. One of my co-workers was working on a change that got approved he had to change the IP of a dns entry pointing to a load-balancer. The change involved inputting the IP address of the DNS server and the load-balancer's IP as well. Needless to say he mixed up the two IPs causing a several hour outage. The moral of the story is that you can architect the shit our of everything try to avoid these problems but there is always going to be a case that you did not cover and it can cause serious outages just like this one.
- wheelinsupial 4y agoIf manual entries are critical to a system, you can have one person enter the info and someone else to confirm the entry is correct. Not sure how common this is in other industries, but this was the method used at an extrusion manufacturing plant I once worked at.
- thrashh 4y agoYea that’s what we call bureaucracy I joke but the right amount of bureaucracy is good, but people are very careful about adding more of it because you can’t undo it
- azherebtsov 4y agoIt is very common in aviation. That is why at any point in time one pilot is flying the plane and another one is monitoring what the former is doing. The pilot monitoring is responsible for verifying the actions made by pilot flying not to punish or blame anyone, but just to make sure that intended action (pronounced by pilot flying) is executed correctly. It is possible that, for instance, you wanted to set value to be X but accidentally rotated a switch a little more. There are also other responsibilities of course. But, yes, making action in pairs is something critical in aviation. That does not happen often when it comes to ground software though. You need to remember that NOTAMs are user-generated content, many systems which allows pushing data to these systems are older than me. Nevertheless database maintenance is possible to do in a safe way. It’s not a DNS nor BGP - it’s high level software where tampering files is absolutely avoidable.
- malux85 4y agoDid you read the article? It explains that procedures were circumvented and something in production was edited - something that was not allowed. You can put as many procedures in place as you want, but unless the humans follow these procedures (I.e. not circumvent them deliberately) then someone always has some level of access where they can cause mayhem. Sure you can programmatically enforce this to a large extent, but ultimately there are always some humans with prod access, or BGP editing access, or firewall write access, or access to that validation logic itself! And it requires that the humans follow the organisational procedures in place. From what the article says, it sounds like the humans deliberately circumvented procedures, in which case, it is the engineers fault, and they should be disciplined.
- guitarbill 4y agoWe don't know the "procedures". They could be braindead, overly complex, and/or impossible to follow.
- landemva 4y agoProcedures can be read by filing a FOIA request. I may be interested in reading it.
- CoastalCoder 4y agoSometimes employees are pressured / instructed / trained to deviate from written instructions. It can be justified as "the instructions are out of date", "this is an urgent change", "just get it done!", etc. Heck, sometimes following the written rules is derided as "working to rule" or "takes longer than others to complete the task". I'm just saying that blame for these issues only sometimes belongs on the line workers.
- carlmr 4y agoAnother one: sometimes, especially in large organizations, there are so many rules that you can't find anyone who actually knows them. Making it unlikely they're being followed. In large org thinking, if something bad happens, and you add a rule, then it's fixed. The problem is that this is not compatible with human psychology. This leads to many rules, the rules are unstructured and impossible to learn, the rules probably also contradict each other. In my experience, a rule without automatic enforcement should be the absolute last thing to depend on. If you do this, your org is broken. The Toyota Production System with it's blame on process design instead of humans is still the best way to avoid these issues IMO. Do 5-whys where human error is not allowed as a root cause. Design your process with poka yoke in mind. If somebody can forget something they will. Don't depend on it.
- tyingq 4y agoThe FAA earlier called it a "damaged database file". I'm not convinced it was data entry exactly. For example, running "/some/script > data-file" instead of "/some/script < data file" still fits the very generic wording the FAA is sharing publicly. Could have been accidentally overwriting some idx/dat pair, or some raw database file, etc. Though, yeah, there should still be some controls, troubleshooting tools, logging, descriptive errors, etc, that would have made things more clear and recoverable.
- jeffbee 4y agoAn earlier report said that the operator had unintentionally clobbered a database file (the newspaper did not use the word "clobbered" but it is the correct term of art).
- JaimeThompson 4y ago>At a private company a failure like this would see the CTO immediately fired and even the CEO having to beg the board for forgiveness and keeping their job. If that sort of thing actually happened a lot more C level people would have been fired for over-hiring, but that rarely happens.
- bigbillheck 4y ago> At a private company a failure like this would see the CTO immediately fired and even the CEO having to beg the board for forgiveness and keeping their job. I don't think this is a thing that actually happens.
- briandear 4y agoSee SWA for agreement with your point.
- atonse 4y agoI don't know if this is true of the CTO and senior tech staff but SWA's CEO had just been in there for 8 months, so I wouldn't expect anything to happen with him. But the previous CEO looks like he was the spreadsheet type and rejected these kinds of ideas. So who is to blame?
- kbutler 4y agoIt's a classic system design failure, but it's 75 years old! It predates much of the learning on designing for human failure. (Various elements and automations of the system have been created and updated, so these could have included improvements in design.)
- Nevermark 4y agoWell put. > […] a data file was damaged as a result of a failure to follow government procedures, […] The more critical processes humans are expected to perform, the worse the expected results. That goes 10x for outside or temporary “contractors” who are not going to be experts in a systems unique, esoteric, or unexpectedly irresponsible points of fragility
- thrashh 4y agoBruh you’re complaining about a system that had about perfect uptime for 30 years. It’s only newsworthy because it was so perfect Second, nothing would happen at a private company because they would involve marketing and marketing wouldn’t blame anyone in a press release. The only difference with FAA is that they didn’t let a marketing department write their release. But really the real issue here is that everyone is posting to make wide generalizations from a one time event.
- bsg75 4y ago> At a private company a failure like this would see the CTO immediately fired and even the CEO having to beg the board for forgiveness and keeping their job. Can you share an example of where that has actually happened?
- indymike 4y ago> CTO immediately fired and even the CEO having to beg the board for forgiveness and keeping their job. I doubt that. More likely: the defect would be identified and fixed, someone would start working on a training-wheels editor for the file to ensure the problem doesn't happen again. Any change in revenues or profits would be used by the CEO to drive change in whatever thing needs changing, and the whole incident would be forgotten. Source: have owned four private companies. If everyone is fired immediately, no one learns a lesson and everyone is terrified to give leaders bad news.
- spfzero 4y agoThe article did not say how the bad data was introduced. There is plenty of room for contractor malfeasance or incompetence to be a factor. Working around the safety measures in place, for example. Using an import tool improperly. Directly modifying a database through console, or, custom application software written without customer knowledge. Abusing privileged accounts in non-authorized ways. I'm sure there are many more. The engineering precautions in system design always involve trade-offs. There is no perfect system that could be created for acceptable costs. Ultimately it is always going to come down to humans following the rules, hopefully as few of those very-trustworthy humans as possible, though that's where the FAA may have been lax.
- albertopv 4y agoI have seen CTOs surviving theft of millions of confidential documents of the company.
- briffle 4y agoHey now, I received my $5.21 check from equifax for their huge data breach! I’m sure that taught them a lesson, but I don’t think it’s the lesson I’m hoping for….
- neuronexmachina 4y ago> The failure was in a technical organization architecting a solution that allowed a simple human error in data entry to take down the entire system. This is classic system design and process failure. Bad input should have been caught before it got pushed to prod, before it got pushed to backup databases, etc. There was clearly no testing, no change process, no backup verification, etc. Sec. Buttigieg apparently agrees: https://www.nbcnews.com/news/us-news/software-blamed-faa-outage-three-decades-old-years-upgrade-official-sa-rcna65562 https://www.nbcnews.com/news/us-news/software-blamed-faa-out... > Transportation Secretary Pete Buttigieg told NBC News that he has asked the FAA, "to make sure that there are enough safeguards built into the system that this level of disruption can't happen because of an individual person’s decision or action or mistake."