8 ms·
Testing on production
- dijksterhuis 3y agothis made me chuckle > If GitHub makes a mistake it can affect thousands of businesses but they’ll likely shrug and their DevOps team will just post “GitHub is down, nothing we can do” on some Slack channel. Gonna try and read the rest of this on the lunch break as was surprisingly meaty for a clickbait title ;)
- kubanczyk 3y agoI love the style: > That’s a terrible mistake and in the long run will be the cause of cost overruns, unmet deadlines, increased churn and overall bad vibes. And nobody wants bad vibes.
- trollied 3y ago"Everybody has a testing environment. Some people are lucky enough enough to have a totally separate environment to run production in"
- deleted 3y ago[deleted]
- chpmrc 3y agoHa! Love it.
- ethbr1 3y agoI code at the interface between ops teams (on the business side) of companies and dev teams (on the IT side). One of the things I've realized is that in most unregulated companies (read: non-healthcare/financial) the business side of the house is used to having little or no lower lifecycle. If they want to make a process change, they make it on production work. Granted, they have change control approvals, etc. etc., but the whole dev-test-prod cycle looks extremely different for them, because you can't do certain things without lower environments.
- JohnBooty 3y agoThis hasn't been my experience. I think it depends on how business-critical the application is. I worked at a home remodeling company. Revenue was several million dollars a day. App handled sales, scheduling, logistics, everything. Breaking production was a big deal, it cost us millions per day and created logjams. I would think that most online applications are the same. Even if a simple online web shop goes down you are costing money. What kinds of experiences have you had where testing in production was the norm? because you can't do certain things without lower environments. I agree that this is something many shops REALLY struggle with. One of the most challenging things is exporting or creating some kind of realistic data set for local development use. I think 99% of companies struggle with this.
- wheelerof4te 3y agoThat should honestly be the norm at any larger company. Even at startups, the added initial costs yield more long term benefits with higher-quality products.
- dfox 3y agoWell, the lucky part is the important one. What it boils down is that such system is either self-contained or has rigidly defined outside interfaces. For anything that deals with a physical reality outside of the pure computational realm this tends to be impossible. You are not going to build an entire warehouse to serve as the physical part of testing environment and even if you did so it will not be really useful, because the thing will be different than the production one due to who knows what tolerances involved in building a physical things. In same vein if you interact with external services you can either mock them or use whatever testing environment the communication partner provides, in both cases it is bound to not behave the same way as the actual production environment.
- anoy8888 3y agoI don’t understand why people redefine words just make their point. It can be confusing at best and at worst change the me meaning of words when it becomes viral. “Smart” people means smart people. It shouldn’t be used to mean junior dev who are trying to hard to prove themselves and over engineer or choose the wrong approach. So many words have changed their original meaning because someone decides to write a viral post and redefine words to make a point
- deleted 3y ago[deleted]
- chpmrc 3y agoOh what an accomplishment it would be, to be able to change the meaning of the word "smart" with a single article! (Don't take it too seriously, like I said this is mostly a brain dump, I'm sure there's a lot of stuff that can be improved)
- JohnBooty 3y agoI like your usage of "smart" in the article. I see this challenge a lot in the industry. The young engineers truly are smart, even brilliant, but lack wisdom and experience.
- chpmrc 3y agoI completely agree. "Smart" isn't used sarcastically here. It's an adjective that most young devs would (rightfully) like to be referred to as. But I see experienced devs as less interested in looking/being "smart" (or clever or whatever word you want to use) and just getting things done in a way that allows the org to make money and get rid of BS (unrelated to the above) as much as possible. Maybe there's a better way to outline this difference.
- routerl 3y agoOh right, the "original meaning" of "smart"... so you must mean "pain or ache"? I really don't see how that's relevant to the article. Words change, they always have, they always will. Get over it. And anyway, the article's usage is consistent with the well-established phrase "smart guy", within which the word "smart" carries a sarcastic and derisive tone.
- bhaney 3y agoLove this article. So many great points that I deeply agree with but have never really put into words, and all written in such an engaging style.
- chpmrc 3y agoThank you so much!
- therealchiko 3y ago> The TL;DR is that some (“best”) practices are contextual and understanding when to use them is ultimately what gives us the title of “engineers”. So well put, just today I implemented a feature and kept asking myself if i should be extending the component (leaning more towards OOP) or just add an additional argument to said component. The latter would have stuck more with the current style but I also realized there's no obvious better way, extending made sense and I realized the importance of understanding the nuance and standing up for those design decisions is what I am here to do :) thank for putting that in less words
- NewEntryHN 3y agoGood article, but it's a bit binary on the notion of incident. For the same company, it can be very serious to have a global 1h outage, but not so serious to have the internal admin interface down for 1h. This allows for more fine-grained assessment of the validation required to push to prod: the "checks" only have to test the critical part of the application. Dev exp start deteriorating when the non-critical parts are over-tested.
- chpmrc 3y agoYes criticality is multidimensional, this was a simplification for the sake of brevity. Will add a note. Thank you!
- wheelerof4te 3y agoJust keep the enironments separate, but similar. What works in the test environment, should work in production. Of course, there are always exceptions to this rule. Adapt and modify the code as needed. We keep three environments at work: Dev, Test and Prod. However, dev environments are sometimes neglected and some features land in Test only. So, use Dev as a development playground. Use Test to test the changes made in Dev. If the change is approved in Test, it will go in Prod environment.
- contravariant 3y agoAn interesting perspective I once heard from an information security expert is that there's a difference between risks and 'things that can go wrong'. Something is only an actual risk if it hurts the bottom-line. In particular quite a few things that can go wrong don't carry that much risk, and conversely something that is hard but not impossible to go wrong may carry huge amounts of risk. The trick with this perspective is that after identifying the real risks you can then link the risks and possible mitigations by looking at all 'things' and identifying the ways in which they might fail (and how this may be prevented from happening). This way you can easily identify which mitigations are helping prevent risks and which risks are not sufficiently mitigated. It's a fair bit of work, but it's not complicated and often gives useful insights. What this article basically does is note that you should first asses what risks a failed deployment has, and correctly states that in quite a few cases this risk is low and therefore the mitigations (of which there can be many) may not be necessary and may in fact be doing harm without actually sufficiently preventing any risk.
- Narann 3y ago> something that is hard but not impossible to go wrong may carry huge amounts of risk. I think it's the definition of the black swan theory[1]. [1]: https://en.wikipedia.org/wiki/Black_swan_theory https://en.wikipedia.org/wiki/Black_swan_theory
- chpmrc 3y agoFantastic take. Thank you.
- HL33tibCe7 3y agoRisk = likelihood * severity
- hef19898 3y agoSometimes one has to include detectability as well.
- datadrivenangel 3y ago
- civilized 3y agoI have a dumb question as a non-SWE who is curious about software engineering. I've heard "feature flags" are popular these days, and I understand that that's where you commit code for a new way of doing things but hide it behind a flag so you don't have to turn it on right away. Now, if I want to test in prod, couldn't I just make the flag for my new feature turn on if I log in on a special developer test account? And if everything goes well, I change the condition to apply to everyone?
- sidlls 3y agoFeature flags are just code, like the rest of the software. You can program any feature with it, including auto-enabling it given appropriate circumstances (e.g., the user is logged in to a developer account). Of course, that doesn't work for features available without requiring an account.
- mwint 3y agoFeature flags sound great, but a company I’ve been consulting for has been using them to their own detriment. Seems like many bugs are due to a (production!) user not having the right combinations of flags enabled. There ends up being code to deal with what happens when various combinations of flags are on/off, and that code doesn’t get tested much. And teams spend a lot of time just removing flags. This isn’t a safety-critical app - I really think they’d do better dropping the flags, and just deploying what they want when it’s ready.
- xnorswap 3y agoI'm going to go further and say that Feature Flags are a nightmare and should be avoided. Because instead of just being used to stage roll-out, they get used to configure different environments for different customers. You not only waste time with "Remove feature flag X" stories if all customers end up with the feature, you also slow down the response time of some categories of bugs, because you end up having to stop and check the combination of feature flags to reproduce a bug. And if you end up with a feature that isn't popular except by one customer, not only are you now stuck supporting "Legacy feature Y", you're actually stuck supporting, "Optional legacy feature Y" which is worse. Maybe I'm ranting about "misuse of feature flags", but I don't like to pontificate about how things ought to be, but how in my experience they actually are.
- postalrat 3y agoWe all test in production but some people are in denial and refuse to accept it.
- justincredible 3y ago[dead]
- Shrezzing 3y agoI enjoyed the entire article except this part: > Unfortunately there is no easy way to distinguish between people who are good and need a paycheck from people who just need a paycheck. But you sure as hell don’t want the latter in your team. If you can't tell them apart, then the distinction is unimportant. So if among the group of people who need paychecks, good is indistinguishable from non-good, the comment serves no purpose other than needless elitism.
- solatic 3y agoIt's implied that it isn't easy to distinguish them during interviews. After they join your team, it's very easy to distinguish them.
- Shrezzing 3y agoI've re-read the developer experience section, and I can't see where that implication is established. In that context, the paragraph stands out as an abrupt diversion from the main theme of the section, and undermines the argument of the entire piece. The section defines developer dissonance, and asserts that it's possible to overcome it with reasoned and sensible questioning. If it's possible to overcome dissonance with reasoned questioning, a hiring interview should a prime opportunity to roll out some reasoned questions and head-off dissonance before it enters the organisation in the first place.
- chpmrc 3y agoI can't see how making that statement undermines the rest of the argument. It would help if you could clarify that relationship. And I'm not sure I understand why, if you can't distinguish them, the distinction is unimportant. It's hard to distinguish an edible mushroom from a poisonous one and yet making that distinction makes a huge difference. Interviews are definitely a limited tool to do so btw, this is only something that you realize over time. It's also very easy to play an interviewer if the interviewee's soft skills are better than the interviewer's (which happens often in this industry).
- 3y ago
- chpmrc 3y agoWelp! For some reason someone at HN decided to change the title and bump this down to the 11th position (atm). Not sure what I did wrong here but it feels pretty crappy... @dang any chance you could help here? :(
- deleted 3y ago[deleted]
- chpmrc 3y agoFor the sake of transparency it was explained to me that the title was too "link baity" and that the comment section was a bit too heated. I appreciate the explanation and I agree this kind of moderation is, unfortunately, required to keep things civil and constructive.
- al_be_back 3y ago>> If Tesla makes a mistake in their autopilot software, people might die. In this case, a good "Testing on Production" rule would be to not let customers test your software, period. There's plenty of land and resources to construct towns and cities that simulate real-life commute very accurately. In the case of self-driving (or even autopilot), you're not really testing a feature, you're researching a new product, they difference is vast.
- hulitu 3y ago> Shipping confidence We can define “shipping confidence” as the feeling a mentally sane developer has when they know their code is about to be deployed to production (whether it can be updated over the air or not). A bug which must be fixed in production is much more expensive than a bug fixed during development. People here complain when you bash Microsoft, but their phylosophy was (and still is) let the users test the product.
- tempodox 3y agoEverybody has got a test environment. Some also have a production environment.
- bornfreddy 3y agoDouble negation is hard... :) (yes and no should be switched) > Ask yourself a question: do you have any reason to think that your engineers will not do a good job? If the answer is no: why are they still there? If the answer is yes: let them do their damn job.