15 ms·
What is negative engineering?
- jonathaneunice 4y ago"Negative engineering" as described here is also especially challenging to test. Requires extraordinary backflips to construct reasonable approximations to the edge conditions, corner cases, external factors, and unreasonable contingencies against which you're trying to defend. The same "Happy Path is just 20% of code" imbalance extends through test rig, tribal knowledge, and effort required. In some cases, 20% seems fantastically generous. IME preventing failures that only occur at heavy load or during other errors may be 95%+ of your effort and represent 99%+ of your "value add."
- sublinear 4y ago> They _insure_ the outcomes that positive engineering tools are used to achieve. Ugh such a tacky abuse of words. Context around that sentence implies it should be "ensure". No, this is just yet another ad promoting gimmicky work for consultants to sell. I understand HN is definitely the place to post this stuff, but I was expecting something more interesting.
- PeterisP 4y agoThe context around that sentence does quite explicitly intends to insure the outcomes because you can't really ensure it. I mean, the paragraph literally starts with "Insurance as code”.
- helmholtz 4y agoYou insure _against_ something. You ensure something. GP is right.
- strken 4y agoIf you're talking about the grammar, then the phrase "My car is insured" is fine. You aren't required to specify what it's insured against; that can be left implicit. If you're talking about the language, insurance is wrong: retries do ensure positive outcomes. However, assurance is wrong as well: unlike retries, observability tools don't ensure positive outcomes, they just mitigate the negative ones. It's just bad writing, using a tortured metaphor to explain a definition that doesn't make any sense.
- giantg2 4y agoYou also drink ensure, if you're old enough.
- OJFord 4y agoI think it's emphasised exactly to avoid it being misread as 'ensure', I don't read it as an intended pun. Doesn't really make sense to 'ensure the outcomes', you 'ensure a good outcome' or something.
- Oxidation 4y agoSee, treating this as an add-on is why software has a reputation as "not real engineering". Engineering is all about turning science into products. Anyone with a chemical lab and enough post-grad grunt can make a compound. That's just the start. Getting it to customers, at scale, with a saleable price and not selling poison, contamination, or thin air and not turning your plant into a crater or superfund site is the task. It just happens to involve a bit of chemistry along the way. You can't just complain from the smoking ruins of your startup "but the happy path we tweeted about when we did it in a flask worked fine, how could we know it would be hard to make 50 tons of FOOF a day?" It's not always glamorous, it's not always fun, it's not even always innovative, it's almost certainly never easy, but it is what you're actually selling, and it's why you're the one selling it and not someone else. Now, can you cut corners and make it work, say, 99% of the time and that's good enough to ship it and sit back[1]? Sure you can. But if someone comes along at 99.9% and the same price, or even a higher price and your customers will pay for that extra 9, you're in trouble. [1]: or ship 99% code thinking it's 100% because you didn't check
- runnerup 4y ago> Anyone with a chemical lab...not turning your plant into a crater or superfund site...how could we know it would be hard to make 50 tons of FOOF a day? For additional context to the mental imagery, I thought it would be appropriate to link FOOF: https://en.wikipedia.org/wiki/Dioxygen_difluoride https://en.wikipedia.org/wiki/Dioxygen_difluoride
- wantoncl 4y agoMore imagery: https://www.science.org/content/blog-post/things-i-won-t-work-dioxygen-difluoride https://www.science.org/content/blog-post/things-i-won-t-wor...
- cratermoon 4y ago"Satan's Kimchi" is the name of my new KPop/Black Metal crossover band
- 4y ago
- deleted 4y ago[deleted]
- MPSimmons 4y ago>Negative engineering is the time-consuming and sometimes frustrating work that engineers undertake to ensure the success of their primary objectives You don't need another name for this. This is part of the engineering process.
- angarg12 4y agoGlad to see I'm not the only one baffled by this article. If you've ever developed and maintained software in production, you had to deal with failures. Is just part of the regular software development lifecycle.
- throwaway5959 4y ago
- P5fRxh5kUvp2th 4y agoHow can you prove your blogging chops if you don't name stuff? You see this a lot in academia as well, they absolutely love to name things, and think the teaching and learning of those names is important in some manner.
- croo 4y agoI usuall call this "meta-programming" or "meta-work". The work you need to do in order to make way for actual work.
- drewcoo 4y ago> What is Negative Engineering? "Negative engineering" is a term coined and championed by Jeremiah Lowin, apparently. Most search engine hits on the term are his or kindly refer to him. He uses it to mean coding for the negative test cases.
- bestcoder69 4y agoFor some context in case people don’t know, this is Future.com, the a16z PR site that they’re shutting down soon for lack of interest.
- extasia 4y agoIsn't what the article describes the concept of "defensive programming"?
- dejj 4y ago“Defensive programming” seems more apt to presentation. I also consider it part of mitigation in FMEA. https://en.wikipedia.org/wiki/Failure_mode_and_effects_analysis https://en.wikipedia.org/wiki/Failure_mode_and_effects_analy...
- jasode 4y agoFyi ... other related terminology ... "chaos monkeys" engineering from Netflix, designing for "fault tolerance" , etc. See list from : https://en.wikipedia.org/wiki/Chaos_engineering#See_also https://en.wikipedia.org/wiki/Chaos_engineering#See_also Another list is "cross-cutting concerns" which can include error-detection & correction, process monitoring, etc: https://en.wikipedia.org/wiki/Cross-cutting_concern#Examples https://en.wikipedia.org/wiki/Cross-cutting_concern#Examples And "site reliability engineering" popularized by Google: https://en.wikipedia.org/wiki/Site_reliability_engineering https://en.wikipedia.org/wiki/Site_reliability_engineering
- giantg2 4y agoI'm a negative engineer.
- mdip 4y agoAs I was reading this article, that itch in the back of my head kept being triggered. I didn't encounter anything obvious, but I felt like something was missing. Then I remembered a project I inherited at a past job. tl;dr Good idea, but it can be easy to end up with very complex code flow if best practices/language-specific patterns aren't followed. I'm not sure if my experience is an example of Negative Engineering success or failure (a bit of both, perhaps). I had inherited a tool that monitored a conferencing environment looking for specific conferences that also were not started and had non-moderators waiting to be admitted[0]. It was a code base written by someone who was not familiar with the language it was written in, resulting in code that was very non-ideal. I had joked that the fact that the system worked was "an accident". It was a perfect example of "failing it's way to success". Every time a targeted conference was discovered, an exception was thrown. When the exception was thrown, the catch block would retry the request slightly differently. It would then add an invisible (which started the conference), monitor for the conference state to indicate it was started and remove the invisible user from the conference. The add/remove had to be done because the "invisible users" could not join two conferences, so we had a pool that would grow/shrink as needed. Nearly every step of the way, of course, there were many circumstances that were appropriate to retry, to catch and recover and (ultimately) to crash[1]. Now, my approach to this kind of a problem is "if you can check that a request will fail and fix it before making it, do that", so I have the "invisible user pool" ensure it adds users before the pool is empty (and if it is empty, creates/returns a fresh user instead of throwing). Nearly every command that could throw had an invariant that could be confirmed before the call yet the developer opted to let things crash and handle the problems as they surfaced. You can imagine the kind of Rube Goldberg device each part of this solution was. Part of that was the language and the developer's inexperience with it. Abusing Try/Catch, as this developer did, made the code feel like it was just a bunch of GOTO statements. The fixed solution had nearly as many Try/Catch blocks -- it's a distributed network application, after all, nearly every call had something external that could go awry. The difference was that there was an obvious "happy path" and the "catch" logic had been limited to either "sleep/retry", "abandon" or "fix; returning to the happy path". On the balance, my experience would suggest that Negative Engineering, when applied exclusively ... works[2]. You just might not be able to explain why it works very easily if it's designed carelessly and (at least in some languages) the patterns you're required to use lend themselves to code that can be challenging to follow. [0] Not a terribly difficult thing to do but not something that the API provided a way to solve so it involved listening for conferences with the API and reading data from the reverse-engineered database the service used... not a difficult task. [1] The service had its own supervisor which handled recovery so for application failures (database unreachable/conferencing environment down), the right thing to do was exit. [2] It logged critical errors left and right, but I wasn't put on the project to fix bugs, I was put on it to add a feature ... something that could no longer realistically be done due to the tech debt/complexity the solution had reached being designed the way it was.
- cratermoon 4y agoFor the naysayers, yes, this should be part of the engineering process. But there are thousands of teams of programmers not doing this. Arguably, they aren't doing engineering, they are coding, or something else. In which case you can correctly say "this is be part of the engineering process", and point out that teams (some of which I've consulted with) who ignore this aren't doing engineering. And yet, the number of systems embedded in our daily lives that are built without "the engineering process" is stunning. They are not typically the essential systems, but just everyday things we depend on and for which the loss would be much more than inconvenient, but short of catastrophic. The author's points may be a rehash of old news, or even trivial, for a substantial portion of HN readers, but for many programmers and their managers, this could make them part of the "lucky 10,000"[1]. 1 https://xkcd.com/1053/ https://xkcd.com/1053/
- kshay 4y ago> Imagine how frustrated you’d be if every now and then, an application simply refused to add items to your cart, navigate to a certain page, or charge your credit card Sort of interesting that this is framed as a wouldn’t-it-be-awful-if hypothetical; I feel like on the contrary, and for some value of “every now and then,” this is most people’s daily experience of interacting with most commerce websites. > The truth is, these minor refusals happen surprisingly often, but users never know because of systems dedicated to intercepting those errors and running the erroneous code again. ...Or users do know, and just try again? Fault tolerance has been outsourced to the customer.
- twobitshifter 4y agoHad the not charge your credit card issue happen the other day. I had to replace my credit card and the power bill was set up to autopay on the old card. I logged in and entered the new number, an unexpected error occurred, I enter it and try again and again. I then notice at the top the account says “power” and there’s nothing due. Clicking switch account gives me the option of “home”, which has the past due balance. What had happened is that my rental from two years ago was still in their system and was the default account for payments. Rather than having one way to pay the bills you have a modal interface to make payments. When I switched to my house, the CC went through no problem.
- aatd86 4y agoAnd yet many programmers do not like reasoning about errors because it "hInDErS tHe hApPy pAtH"... :) That's funny to me. Actually, I don't even know why call it negative engineering. That should be engineering, period.
- zmgsabst 4y agoA quip I picked up from a Boeing engineer: “Engineering is designing failure.” Obviously that has darker tones in aerospace, but I think it applies to even CRUD apps: - my app will choke and fall over if TPS exceeds 1k/sec - my app will drop data when we exceed 2TiB in the DB total - etc. You’re “engineering” when you design and control those modes in your system.
- airbreather 4y agoThe dangers of implicit state https://medium.com/@DavidKPiano/the-facetime-bug-and-the-dangers-of-implicit-state-machines-a5f0f61bdaa2 https://medium.com/@DavidKPiano/the-facetime-bug-and-the-dan...
- deleted 4y ago[deleted]
- juancn 4y agoThat's just sound engineering. Besides the usual quality considerations (correct, available, performant, etc.), well engineered software systems should be: - observable: you need to be able to tell the state of the system to a sufficient degree (metrics, status reports, etc.) - operable: it should be possible to influence the behavior of the system even in case of failure (feature flags, tools for retrying or pausing jobs, throttle processing speed, etc.) Failure modes should be considered as early as you can, and provide sufficient tooling to deal with them. Most of the effort on large scale or mission critical systems is dealing with failure modes, rather than the happy path and making sure it's extremely difficult to make it fail catastrophically without recourse. Failure should be considered business as usual and dealt with accordingly, hopefully automatically but you need to leave tooling in place for those situations you didn't consider.
- PhasmaFelis 4y ago> Imagine how frustrated you’d be if every now and then, an application simply refused to add items to your cart, navigate to a certain page, or charge your credit card. The truth is, these minor refusals happen surprisingly often, but users never know because of systems dedicated to intercepting those errors and running the erroneous code again. I want to live in the world this guy is from.
- rightbyte 4y agoI've always assumed it was the internet connection or magnetic strip or chip on the credit card that was dirty when terminals failed but were happy on a retry. Maybe it was bugs all along?