20 ms·
Normalization of Deviance (2015)
- calmdown13 4y agoThis rang so true for me. I’m constantly rediscovering things that I already understood well at my previous job. Once you get acclimatised to the current system, so many previously obvious learnings fall by the wayside.
- PaulHoule 4y agoI like the bit about Let's look at how the first one of these, “pay attention to weak signals”, interacts with a single example, the “WTF WTF WTF” a new person gives off when the join the company. and kinda wonder if a company that prioritized not getting this reaction from new hires might find it is the most impactful thing they can do in terms of culture.
- chaps 4y agoBeen told that before. When I spoke up, I was told that I was new and shouldn't talk about things I know nothing about.
- PaulHoule 4y agoIt's not (1) "reacting to the reaction" which is the endpoint but (2) not having that reaction. If (1) is important it as because that is the path to (2). I'd say a company that has accomplished (2) has cut the workload in hiring employees by 30-50% in the sense that every employee who has reaction (1) either internally or externally is at risk for being disengaged or leaving soon. Not only that but you are probably wasting your dev's time and could get dramatically more productivity out of them if you aren't WTFing them to death.
- giobox 4y agoThere's often a wise tradeoff between criticizing systems you've just seen after being at the company 5 minutes and actually spending some time at the company to learn the historical context of why the thing you think is insane/shit is insane/shit before telling everyone who built it how insane/shit it is. People generally don't wake up in the morning and go into work motivated to make insane/shit things - context, tech debt and business realities all mount up and even the best of us can end up making choices that in isolation look crazy. There are of course companies who are really bad and you may well be right, but so many times I have seen in my career a young new hire storm in and think everything is shit without paying heed to the context and historical pressures. The best thing you can do in many cases is spend the first ~six months at a new tech company trying to understand that context, and indeed I think more mature engineers generally do.
- chaps 4y agoYou (in fashion of the article) missed the point I made: I was explicitly asked to speak up as the new employee and when I did I was told to stop speaking up. When I brought it up in the meeting I gave my two week notice, he admitted to saying that and apologized.
- dec0dedab0de 4y agoThat's some bullshit. you did the right thing. When a new employee joins and notices that something sucks the answer should be something like these: We know, but haven't had time to fix it, maybe we'll assign that to you when you're caught up. We didn't think of it that way, good catch, lets go into detail later. Yeah, but doing it this way makes this other thing easier, we'll show you that when you're ready. or even: I don't know, my brain is fried with this project, can you ask again in a few months?
- a4isms 4y ago"Making good decisions, therefore, requires understanding past decisions. Without knowing how things came to be, it’s easy to make things worse." —https://thoughtbot.com/blog/chestertons-fence https://thoughtbot.com/blog/chestertons-fence
- chaps 4y agoI was once told this after raising the issue that "hey maybe an API that responds with root ssh passwords is a bad idea, and our clients are going to be pissed once they find out." And.. I was right. So often, citing Chesterton's fence is significantly more naive than what it attempts to criticize.
- justin_oaks 4y agoI agree. Chesterton's fence doesn't mean that if you don't know why the fence is there, don't ever move it, under any conditions. It means try to find out why it's there before moving it. In many cases in my career, I've seen code that doesn't make sense or seems like a bad idea. The person who could explain why it's there has long left the company. Am I afraid and leave the screwy stuff there, while citing Chesteron's fence? Hell no. I'll change it to do the right thing. This results in either exposing the reason why it's there, or showing that it really was unnecessary/bad. If something breaks from the change then it's good that I can finally document what wasn't documented before. So either way it's a win.
- AnIdiotOnTheNet 4y agoTo be fair, you really shouldn't. You know nothing of the constraints that people are operating under, or the political or cultural landscape you're dealing with, so you just come off like a preachy academic.
- chaps 4y agoSee the other comment I made: I was explicitly asked to speak up.
- AnIdiotOnTheNet 4y agoIf that's the case, then yeah fair enough. If someone doesn't want your opinion they shouldn't ask for it.
- lo_zamoyski 4y agoThat doesn't seem like a healthy standard b/c it grounds decision making in appearances rather than principles and prudential judgements. Certainly, such feedback or opinions can be worth considering as a way of getting at what principles are being violated and deciding whether these violations are tolerable or what ought to be done about them. A fresh pair of eyes could help. But an untrained pair of eyes might also not be qualified to discern the right course of action.
- PaulHoule 4y agoYes and no. But 2/3 of the time the problem is that the company has no documented build process or something obvious like that . They have time to spend 2 years failing to deliver a product because they don't know how to build it, but when somebody asks "How do we build it?" the answer is "Don't waste our time asking stupid questions." Of course they have been wasting time not knowing how to build the system, the guy who started the project might be able to hit F5 in 15 different windows and get it to sorta kinda work, but new hires they are hiring to work on it are quitting right away and somehow they can never get it into production. There should be no controversy at all that complete instructions for installing everything required for a dev to build the project and work on it should exist and it should be possible to complete this task in hours, not the weeks that it frequently takes. And, no, "docker" is not an answer to this anymore than "The F5 Key is a Build Process" https://blog.codinghorror.com/the-f5-key-is-not-a-build-process/ https://blog.codinghorror.com/the-f5-key-is-not-a-build-proc... It is not "Docker" that solves the problem, it is the discipline of scripting the image build process into a dockerfile. If you know how to write a dockerfile you can write a bash script that runs in 20 seconds as opposed to having Docker spend 20 minutes downloading images and then crash because of a typo. You are right that a company might have good reasons for doing things in an unobvious way, but most of the time when nobody at a company claims to understand what the company is doing except for the CEO and people aren't too sure about the CEO, it is the fault of the company lacking alignment, not a natural property of freshers.
- skissane 4y ago> It is not "Docker" that solves the problem, it is the discipline of scripting the image build process into a dockerfile. If you know how to write a dockerfile you can write a bash script that runs in 20 seconds as opposed to having Docker spend 20 minutes downloading images and then crash because of a typo. The problem is the bash script may end up depending on poorly understood aspects of the local setup (global config files, installed packages, etc) - it might work fine now, but then nobody runs it for 12 months and there’s some churn in personnel and suddenly people are trying to work out why it crashes. Dockerfiles can avoid some of that stuff, although not always (e.g. the common problem that if you don’t fix the versions of packages to be installed, an updated package is released which then breaks the Dockerfile)
- GartzenDeHaes 4y agoIn my experience, you cannot change an organization's culture with rules, mission statements, listing values, or giving speeches. They only way is to take down the old culture bearers. Some of them may be managers, but more often they are employees who have gained some organizational power. You'll often find them at the center of sticky organizational spider webs with approval processes such as purchasing and service administrators. To change the culture, these people have to go. Firing them may not be feasible, but there are other options. Dethroning them in the form of a promotion or even just physically moving them can be effective. When people don't have to jump through their hoops anymore, they lose their organizational power.
- Ensorceled 4y ago> They only way is to take down the old culture bearers. ... To change the culture, these people have to go. I was briefly head of engineering at a company that had several "old culture bearers" that made change impossible. I was something like the 3rd or 4th engineering leader over the space of a year. Apparently the person after me was actually allowed to fire a few of these people and was able to turn things around.
- PaulHoule 4y agoMore than once I had to leave in order for an employer to take action against a problem employee, including my boss once. In these cases the employer suddenly realized that it made no sense to attempt to replace me without removing the reason that drove me out.
- Ensorceled 4y agoThat is, unfortunately, very common. Though sometimes the toxic people drive the company into the ground first.
- jkaptur 4y agoJust to give a different, concrete, perspective (and push a hot button HN issue), I've spent a fair amount of time working on extremely large web applications, and by far the #1 "WTF WTF WTF" thing that new hires say is "what do you mean you aren't using $TODAYS_HOT_JS_FRAMEWORK??" Once you get away from "should we use version control" and into actually difficult software engineering questions, it's not clear how to balance a fresh perspective vs. an experienced (normalized? tainted?) view. I wish the article went into this more. Like, how does the new hire (or anyone else) know the difference between "learning the complexity of the new system" and "internalizing/normalizing the deviance of this culture"?
- aidenn0 4y ago> Like, how does the new hire (or anyone else) know the difference between "learning the complexity of the new system" and "internalizing/normalizing the deviance of this culture"? If a new hire can't checkout, build, and test the software on the first day, then there is likely something either wrong with the hire or the infrastructure. A sufficiently old and arcane software system might take weeks before a new hire can make even a simple change, but that shouldn't impact those three items.
- Jtsummers 4y agoAlong with this "how long to spin up the new hire" issue, one of my first (if not the first) questions when trying to help people improve their processes related to software is: > If the user/client asks you to make a small but not trivial change, how long would it take to update and deploy the program? I have had answers ranging from "A couple hours" to "A year" (yes, they were serious). Most were in the 1-3 month range, though, which is pretty bad for a small change. It also makes it apparent why a bunch of changes get batched together whether reasonable or not. If a single small change, single large change, or collection of variable sized changes all take a few months to happen, might as well batch them all up. It becomes the norm for the team. "Of course it takes 3 months to change the order of items in a menu. Why would it ever be faster than that?"
- 4y ago
- e_i_pi_2 4y agoThis is something we actively try to take advantage of at my company - we know that we've grown comfortable with architecture that may not make a lot of intuitive sense, so when we have new people join we try to make a list of confusing concepts so we can try to clean them up. In the ideal case the new person is able to do the cleanup, so we get a more intuitive design and they learn the surrounding architecture more along the way
- rr888 4y agoThe problem is when you join a high performing team and organization. Is a WTF something they should fix your you need to recalibrate what you think is normal?
- renewiltord 4y agoIn The Field Guide to Human Error Investigations by Sidney Dekker, he quotes someone else saying something like: > Everything that can go wrong will go right. Murphy's Law then manifests from escaping disaster through repeated iterations of taking risks where most things play out well anyway. I have to laugh at the "append z to the end" strat at Google, though. That's a good one.
- chaps 4y agoThese stories ring so, so true. Once worked at a company whose infrastructure issues were so deep and festering that after fighting a fire, my boss told me, "If you go to the press about this, the client will sue us and everyone who works here will lose their jobs."
- aliqot 4y agoThat's what I never understood about this story; did you guys have any suspicion it would dump radiation into the patient all at once, or was this like a concurrency bug
- chaps 4y agoWe weren't able to reliably install security daemons on a client's machine because the entire automation system didn't account for autoscaling. The issues were raised well before I joined and the project head legitimately didn't understand it as a problem that needed solving. The hosts were for a presidential candidate's webserver, and they noticed the webservers were missing security daemons days before the election.
- aliqot 4y agojeez thats a rough spot to be in. did you stick around to fix it or just get the hell out of dodge after that?
- chaps 4y agoI did what I could with a handful of selenium scripts, then hit a road block because we didn't have ssh access to a chunk of the autoscaling hosts. Gave up after that, told the customer rep to tell them we can't do it, and gave my two week notice about a month later.
- aliqot 4y agoOuch, that has to be rough to endure. I'm glad you seem to be in a better place now. Good on you for doing the right thing and getting the hell out of there when your options ran out.
- dec0dedab0de 4y agoI'm still reading, but I just got to the part about flaky, and I got annoyed because there are clear use cases for flaky or pytest-retry. If you have an integration test that relies on an unreliable system you do not control. Sure you can mock it out for a unit test, but if you want to make sure you catch breaking API changes, you need to hit the actual system. And if it works after retrying it a few times, then so be it. no need to throw shade.
- Aeolun 4y agoI think the issue is people using it when they’re too lazy to fix the test case.
- namaria 4y agoI find quite interesting that people will prefer a highly malleable language like Python, and then orgs have to adopt testing to get around all the inconsistencies caused by absent type system. And then people will write libraries to get around the pesky tests to get their flexibility back. It's fascinating really... Complex systems are always in partial failure mode and that applies to collective optimization challenges. Organizations will always be stuck in local optima in most domains.
- dec0dedab0de 4y agoType systems do not replace testing, and if a test works after retrying it then it is probably not something that a type system would be able to catch.
- lmm 4y ago> Type systems do not replace testing They're a good substitute for many of the use cases of testing. > if a test works after retrying it then it is probably not something that a type system would be able to catch. Type systems are pretty good at catching incorrect concurrency logic these days, and getting better all the time.
- stevehawk 4y agoThis is a big term in aviation, because in most cases in order for something catastrophic to happen it requires a lot of things to have failed. And one way to ensure that enough things fail is to start deviating from your maintenance, inspections, or general responsibilities. Related: the swiss cheese models.
- kurthr 4y agoIt's also why things in aviation are so fixed and difficult to change. Not having any new civil aviation planes for 30 years worked... how about 40, 50? When will it break? Well, when someone develops an easy to build, easy to fly, inexpensive experimental craft and zillions of people do it all at once (hasn't happened yet). One of the more interesting things I've found is that a huge number (easily a majority) of instructors are recently trained flyers, because there is a pipeline to train them and they're cheaper than using experienced pilots (esp for multi-engine and more complex airplanes). They also know all the ins-outs of the training and rule books (with recent changes) so they know how to pass all the tests and how to teach that. Sooooo you have a bunch of inexperienced pilots teaching all the new pilots... there's likely a failure there, but it hasn't reared its head. We still have a lot of ex-military folks around who didn't learn that way. Who do you want flying when things go bad? People who have spent many hours with things about to go bad (military, emergency/fire, sail plane pilots) who have experience dealing with it. Those people can also be fun/terrifying to fly with, because they will take risks.
- sokoloff 4y ago> Not having any new civil aviation planes for 30 years worked... The A220, A350, and A380 are all newer than 30 years old. (A321 is barely younger than that and A330 barely older.) Boeing has released the 777 and 787 in the last 30 years. The Cirrus SR20 and SR22 are newer than that, as is the SF50 jet. The Diamond DA40, DA42, and DA62 are newer. The Honda Jet is newer. Cessna has a handful of business jets newer than that. The Embraer Phenom 100 and 300 are newer than that. There are variants of the CRJ newer than that (-700, -900, -1000). That's a lot of new civil aviation aircraft designs in the last 30 years.
- 4y ago
- AlbertCory 4y agoThis guy needs to organize & format his writing better, since he does have really interesting things to say.
- sebstefan 4y agoFirefox has the "reader view" option toggleable with F9 for when you stumble upon unreadable designs, if you want
- jwilk 4y agoIt's F9 on Windows, command-option-r on macOS, and ctrl-alt-r elsewhere. https://support.mozilla.org/en-US/kb/firefox-reader-view-clutter-free-web-pages https://support.mozilla.org/en-US/kb/firefox-reader-view-clu...
- xnorswap 4y agoIf you don't have a reader view in your browser, just paste this into your global CSS: p { line-height: 1.7; max-width: 60em; font-size: 1.2em; margin-left: 5em; } It pretty much fixes the default readability which is essentially zero on this site otherwise.
- sebstefan 4y agoI don't like fucking with the default CSS, my alternative is this bookmark that injects Javascript in the page and temporarily fixes the formatting at one click of a button javascript:(function(){ var bod = document.getElementsByTagName("body")[0]; bod.style.margin = "40px auto"; bod.style.maxWidth = "40vw"; bod.style.lineHeight = "1.6"; bod.style.fontSize = "18px"; bod.style.color = "#444"; bod.style.padding = "0 10px"; })(); And for a dark mode: javascript:(function(){ var body = document.getElementsByTagName("body"); var html = document.getElementsByTagName("html"); var img = document.getElementsByTagName("img"); body[0].style.background = "#131313"; body[0].style.opacity = "1.0"; html[0].style.filter = "brightness(115%) contrast(95%) invert(1) hue-rotate(180deg)"; img[0].style.filter = "contrast(95%) invert(1)"; })();
- 4y ago
- superpope99 4y agoHas Dan Luu ever explained why he doesn't put dates in his blog posts?
- Jtsummers 4y agoMost of his posts seem to be date-independent. To the extent that it matters, you can check the homepage: https://danluu.com/ https://danluu.com/. There you will find month and year of the posts.
- kurthr 4y agoThere's been a lot of work on reliability of complex systems and how they operate. What has been found is that it is almost always necessary to have failure (degraded operation) modes that prevent system failure, and the more complex and more hazardous failure is the more modes develop. In these systems it is found that they are almost always operating (or transitioning between) failure modes. Often multiple operational failure modes are simultaneous. It becomes very important to test the system in each of it's failure modes and their combinations to maintain high up time. https://how.complexsystems.fail/ https://how.complexsystems.fail/ is an example, but there are many. Human work, development, and maintenance is itself a system that interacts with these critical systems. Frankly, failure to fail causes failure (thus chaos monkey). The mythical man month is almost a sub category of these failures as are HR hiring processes and other BS. Being too successful and not having competition (or similarly sclerotic competition) can be as much of a hazard as "move fast, break things".
- euroderf 4y ago"When a fail-safe system fails, it fails by failing to fail-safe."
- kerblang 4y agoGreat stuff - I think this goes in the "required reading" list. The tech industry tends to revolve around "I'm a super-rational robotic genius" thinking that can't accept the existence of its own irrational tendencies, to the point that it becomes ridiculous.
- SideburnsOfDoom 4y agoThe standard "Required reading" text on the subject is "The Field Guide to Understanding 'Human Error'" By Sidney Dekker
- goostavos 4y agoGreat book! (Although, it could have used a more aggressive editor) Reading it felt like a personal attack in many places. However, reading it forever changed how I think about things. It's a much more useful framing for everyone involved if you start with the question of "why did they think this was the right thing to do?" as opposed to "this person made a bad choice / mistake". My (extraordinary) impatience naturally predisposes me towards the latter, but the core argument of the book is that that's lazy -- you can hand wave away anything and everything with "operator error".
- deleted 4y ago[deleted]
- irsagent 4y ago"rictus of horror" - What a set of words to describe a response.
- justin_oaks 4y agoI welcome others to share stories of the normalization of deviance in their companies. One company I worked had no unit tests, no infrastructure as code, and no build server. This held strong for a while until enough developers implemented some unit tests, infrastructure as code (e.g. terraform), and a build server as skunkworks projects. Eventually management tolerated them, but never endorsed them. Some teams at the company still never embraced good practices because it wasn't forced on them. I guess I've never worked at a company that valued unit tests across the whole of the engineering team. I introduced them and implemented them on my own team, but others ignored it.
- ctroein89 4y ago> and no build server Personal experience is that a build server normalizes deviance. "But it works on the build server" we used to say, as, with time, it become harder and harder to build locally. "Just fix your environment!" we used to say, when it was the build system that was actually at fault. "It's all so fragile, just copy what we've done before!" we then said, repeating the mistakes that made the build system so fragile. Eventually, the build system moved into a Docker image, where the smells where contained. But I'm still trying to refactor the build system to a portable, modern alternative. If we hadn't had a build server, we'd have fixed these core issues earlier and wouldn't have built on such a bad foundation. Devs should be building systems that work locally: the heterogeneity forces better error handling, the limited resources forces designing better scaleability, and most importantly, it prevents "but it works on the build server!".
- lmm 4y ago"But it works on my macbook" is even worse than "but it works on the build server". If you have a build server you're at least forced to make sure it builds in two places (your own machine and the build server) before you merge rather than only one.
- MikePlacid 4y ago> “But it works on the build server" we used to say, as, with time, it become harder and harder to build locally This got me puzzled for a couple of minutes. Yeah, that “WTF, WTF” moment. Then I realized that our build “server” comprised of 12 different platforms (luckily reduced to just 6 in the later years), so to pass a build in production was a bit harder than to build locally.
- pizzaknife 4y agoi routunely remind everyone,"we're all mercenaries." i have marginal control over who i manage. The Product isnt saving the world, but it is allowing us to live reasonably and with a clear soul at the end of the sprint. The reason i say the "mercenary" bit is simple: weigh your dreams against blood and gold and compromise.
- martopix 4y ago> Have you ever mentioned something that seems totally normal to you only to be greeted by surprise? Once some (foreigner) person was surprised at my dipping toast with Nutella in my latte. I was equally surprised by his surprise.
- shadytrees 4y agoformatted for wide monitors https://ddanluu.com/wat https://ddanluu.com/wat
- somat 4y agoA thought experiment. When is it "Normalization of Deviance"? and when is it a "Efficiency Optimization"? I mean, the difference is pretty clear after something has failed, But very murky before.
- jpollock 4y agoIt is Efficiency Optimization when you know why the rule is there, and having made an estimation of the risks, perform a cost-benefit analysis. aka "Chesterton's Fence" Otherwise, it's "Normalization of Deviance": * The build is broken again? Force the submit. * Test failing? That's a flaky test, push to prod. * That alert always indicates that vendor X is having trouble, silence it. Those are deviant behaviours, the system is warning you that something is broken. By accepting that the signal/alert is present but uninformative, we train people to ignore them. vs... * The build is always broken - Detect breakage cause and auto rollback, or loosely couple the build so breakages don't propagate. * Low-value test always failing? Delete it/rewrite it. * Alert always firing for vendor X? Slice vendor X out of that alert and give them their own threshold.
- manicennui 4y agoUnfortunately I don't find that most software engineers understand the difference between actually determining costs and benefits and choosing to make certain tradeoffs and rationalizing whatever choice they already made.
- jpollock 4y agoI think that's ok, For me, it's more about "change the system, instead of ignoring it". Once you change the system (document/rules/alerts/etc), then if it breaks, you change it again and learn the lesson. Both are conscious decisions by the org.
- t3estabc 4y ago[dead]
- nilespotter 4y ago[flagged]
- _8j50 4y agoIt is not.click.
- dang 4y agoCan you please not post unsubstantive comments? It looks like you've been doing it repeatedly, and we're trying for something else here. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- nilespotter 4y agoDang, I try to behave, I really do. That little quip so clearly suggested itself it would have been physically uncomfortable not to make it. I will re-read the guidelines today and redouble my efforts.
- Logans_Run 4y agoI'm not sure if this link has already been posted but have a look at How I Almost Destroyed a £50 million War Plane and The Normalisation of Deviance. https://www.fastjetperformance.com/blog/how-i-almost-destroyed-a-50-million-war-plane-when-display-flying-goes-wrong-and-the-normalisation-of-deviance https://www.fastjetperformance.com/blog/how-i-almost-destroy...
- aj7 4y ago“Let's say you notice that your company has a problem that I've heard people at most companies complain about: people get promoted for heroism and putting out fires, not for preventing fires.” My first day at work at big-laser-company. Manufacturing engineer for a laser (then) so complex, it required a PhD to solve problems to get units out the door. The product was a ring laser. What that means is that the laser beam travels around in a race track pattern inside the laser before getting out, not a back-and-forth bouncing between two mirrors. Now this laser could be tuned to any wavelength by suitable setups and machinations, and once there, would “scan” a small amount about this wavelength, enabling scientists to study tiny spectral features in atoms and molecules with great precision. I knew all this shit. I was a Berkeley-trained physicist that built precision lasers out of scrap metal for my thesis. First day of work. I walk into the final test lab. The big laser was happily scanning away. The bright yellow needle-like output beam was permitted to hit the lab wall. As the laser scanned, the beam was MOVING on the wall. Whereupon, first day of work, I exclaimed the most obscene four words in manufacturing, for all to hear, “You can’t ship that!” (“Beam pointing instability” is detrimental to almost any laser application. It turns out that during scanning, an optical element was rotating, on a shaft, inside this laser. This mechanical motion caused beam motion.”) Well, I got an immediate reputation as a negative guy. (You can tell it’s deserved.) The solution was to retrofit 28 lasers in the field, mostly in Europe, with a component that cancelled the movement, on an expensive junket by a service guy. Who was hailed as a ”hero.”
- KMag 4y agoInteresting. I had an internship at a company that did inertial navigation, mostly for defence applications. I only knew of ring lasers for use in gyroscopes. (Send a laser around a loop wave guide/fiberoptic, and any translational acceleration cancels out going out and back, but any acceleration in rotational velocity in the plane of the ring/rotation vector perpendicular to the ring shows up as a Dopler shift. Tune the laser to have a standing wave, and rotational acceleration shifts the nodes of the standing wave around the ring.) I had a colleague who got called up when a Trident missile MIRV bus fell off a forklift and he had to do simulations to tell the Navy if it was still good or needed to be brought back in for rework/recalibration. My understanding is that either the MRIV bus itself or its container has integral devices that record peak 3-axis acceleration for just such a scenario. I imagine they're as simple as a few precise weights on a few wires with precise failure strains, so you can bracket the peak acceleration by which wires broke and which survived. On the one hand, it's great to have more accurate nukes, which allow lower yields, smaller stockpiles, and presumably smaller craters if everything goes sideways. On the other hand, "surgical" nukes result in it more likely that one side will use them and gamble that the other side won't massively retaliate.
- dang 4y agoRelated: Normalization of Deviance (2015) - https://news.ycombinator.com/item?id=22144330 https://news.ycombinator.com/item?id=22144330 - Jan 2020 (43 comments) Normalization of deviance in software: broken practices become standard (2015) - https://news.ycombinator.com/item?id=15835870 https://news.ycombinator.com/item?id=15835870 - Dec 2017 (27 comments) How Completely Messed Up Practices Become Normal - https://news.ycombinator.com/item?id=10811822 https://news.ycombinator.com/item?id=10811822 - Dec 2015 (252 comments) What We Can Learn From Aviation, Civil Engineering, Other Safety-critical Fields - https://news.ycombinator.com/item?id=10806063 https://news.ycombinator.com/item?id=10806063 - Dec 2015 (3 comments)
- csomar 4y agoExtrapolating from these related threads, this article should have 2K+ comments the next year.
- aeturnum 4y agoIf you enjoyed this - I highly recommend watching Adam Curtis' Can't Get You Out of My Head: https://thoughtmaybe.com/cant-get-you-out-of-my-head/ https://thoughtmaybe.com/cant-get-you-out-of-my-head/
- deleted 4y ago[deleted]
- bluedino 4y agoWrite total shit for code, then look like a 'genius' for 'fixing' bugs, only to have them come back again in the future (further looking like a clown to the rest of the team)
- deanCommie 4y agoReality: It is true that EVERY organization is broken in some way or another. You have to find the one that is broken in the way that is tolerable to you. Arguably the closest we know to a panacea in terms of engineering culture and best practices is Google. And what are they now known for? An inability to ship anything meaningful anymore. Spinning around in circles launching and re-launching new chat apps. These are not unrelated. High engineering standards are always in tension with product delivery. As a security engineer once told me, "the most secure system is the one that never gets launched into production." So while Dan is right, and all the examples are right, and things like non-broken builds and a fast CI/CD pipeline are totally achievable, don't learn the WRONG lesson from this which is that when you arrive to a company and notice a bunch of WTFs, the first thing you must do is start fixing them in spite of any old timers who say "Actually that's not as bad as it seems". Sometimes they're wrong. USUALLY, they're right.
- Wistar 4y agoAOPA: Normalization of Deviance in Aviation https://www.aopa.org/news-and-media/all-news/2015/december/07/the-normalization-of-deviance https://www.aopa.org/news-and-media/all-news/2015/december/0...
- theptip 4y agoAs a new hire there is a line to walk between on one hand, using your outside/fresh perspective to provide valuable insight to the org, and on the other, complaining (or appearing to complain) about decisions where you don’t have full context. Many of the examples in the OP are probably closer to the former, but my general advice here is to keep lots of notes about what seems broken, and revisit in a month or two. Sometimes you gained context that explains why something is actually sensible. If it still seems crazy with context, you can now bubble up the feedback with confidence, and also having hopefully built some respect and trust from the team to make the message land better.
- afahad 4y agoI couldn’t finish reading this. Horror film. A screenplay to dystopian dread.
- quickthrower2 4y ago> It's technically possible to use @flaky for that, but in practice it's used to re-run the test multiple times and reports a pass if any of the runs pass This is useful and fine. Someone wrote a test and it now hits a race condition or something and occasionally fails. Let’s assume we are very confident it is problem with the test not the product. Choices: Spend a sprint trying to fix it right now regardless of priority. Turn it off and lose that coverage. Buy some time. In this context it makes sense. As long as their is a procedure to address these in some sane timeframe. Maybe that is an example of normalization of deviance. But I think if it is discusses and trade offs thought through it is an OK thing to do at times. Remember most development is not green field. You inherit a system when you start a job.
- overengineer 4y agoHere are a few examples from real-world history that reflect problems discussed in the article: - The Space Shuttle Challenger disaster in 1986 was caused by the normalization of deviance, where engineers became accustomed to problems with the O-ring seals and began to accept them as normal. This led to the eventual catastrophic failure of the shuttle's launch, killing all seven crew members. - The 2008 financial crisis was caused in part by a normalization of deviance in the banking industry, where risky and complex financial instruments were routinely used without proper oversight or understanding of the potential risks. This led to a widespread collapse of the financial system and a global economic recession. - The Volkswagen emissions scandal in 2015 was caused by a normalization of deviance in the automotive industry, where engineers and executives became accustomed to cheating emissions tests and misleading customers about the true environmental impact of their vehicles. This led to significant financial and reputational damage to the company. - The Theranos scandal in 2018 was caused by a normalization of deviance in the healthcare industry, where the company's leaders became accustomed to misrepresenting the capabilities of their blood testing technology and misleading investors and customers about its accuracy. This led to significant legal and financial repercussions for the company and its executives. (ChatGPT)
- thomastjeffery 4y agoOur society places an enormous value on "competitive obfuscation". As an obvious result, our society does an incredible amount of work maintaining that obfuscation. --- I've heard estimates that 20% (1/5th) of all healthcare-related spending in the US is overhead from insurance determinations, paperwork, etc., and that that 25% (25%/125%=1/5th) extra spending (relative to 100% of the rest of healthcare expenditure) does not exist in single-payer healthcare systems, like those used in Canada, Germany, and every other developed nation in the world. What do we get from that extra spending? What substantive difference does that obfuscation provide? The main difference I see is "explicit opportunity cost". Instead of deciding ahead of time that we will pay for any arbitrary healthcare need (as a single-payer program), the opportunity for each individual healthcare act is given a price, and groups of priced opportunities are provided by subscription-based insurance plans. Every person has to find, apply for, and pay for an insurance plan that will meet their current and future healthcare needs. Because that is explicit, there is leverage available to manipulate each opportunity cost, and even the opportunity of each person to have that opportunity provided to them. So what does that leverage even look like, and who is using it, and for what purpose? Politics. Instead of care being determined by your doctor, access to each type of care is explicitly made available (or unavailable) by your insurance plan. That's a huge attack surface for political motivation. There is currently a dextroamphetamine (Adderall) shortage in the US. The other day, I went to my pharmacy to pick up my prescription for 30 generic Concerta (methylphenidate extended release, another stimulant medication used for ADHD), and learned that all they had left were 16 brand-name Concerta. I was lucky enough to have that covered by my insurance. Many different insurance plans would not have provided me that opportunity. Why is there a shortage? Despite a significant increase in ADHD diagnosis last year, the DEA refused to raise the limit of Adderall that can be legally manufactured. Why? Because there is a long-standing political conflict between stimulant addiction prevention and ADHD treatment, and the DEA is positioned at one side of it. That same political conflict is why some insurance companies have outright refused to include coverage for stimulant medications. Even without a nationwide shortage, some people have found themselves stuck in a position where the opportunity for medication is held just out of reach by the political decision of their insurance company, or the lack of access to insurance at all. The same pattern can be found with practically every type of medical care that is politically controversial: contraceptives, abortions, hormones, etc. Even if you can't get a legislative ban, there is still leverage available to obfuscate opportunity itself. When conservative politicians argue that a single-payer program would be "too socialist for America", the substantive difference they intend to preserve is the political leverage that is baked into the system we have; the political leverage that allows politics to restrict our medical care without a single vote. --- That's just one example. This pattern is everywhere. The only answer is social objectivity. It's a hard problem.