10 ms·
Buried in the "2.11 Frequent rewrites" section, but a great hack for "productivity via a sense of ownership": "In addition, rewriting code is a way of transfer
by Radim 8y ago
Buried in the "2.11 Frequent rewrites" section, but a great hack for "productivity via a sense of ownership":
"In addition, rewriting code is a way of transferring knowledge and a sense of ownership to newer team members. This sense of ownership is crucial for productivity: engineers naturally put more effort into developing features and fixing problems in code that they feel is “theirs”."
- yakshaving_jgt 8y agoThanks for highlighting this. To me it seems an important idea that contradicts conventional wisdom, similar in the way that most people over-encourage DRY, blind to the fact it increases coupling.
- lugg 8y agoI frequently find myself drastically refactoring code to understand it. I don't commit those changes because it's not worth the effort to justify the cleanup to people who treat these rules as gospel. Apparently me spending half a day reading code is no big deal but cleaning it up is a waste of time. Shrug.
- quietbritishjim 8y agoI think the objections to rewriting code are sometimes justified, rather than just blindly following rules. As you said, refactoring code helps you understand it. This means that you can end up feeling like your code is objectively clearer than before you started, but sometimes it's just an illusion caused by the fact that you just (re)wrote it. If there are other people in the company that already understand that functionality, you're disrupting the effort they put into understanding the old version of the code. Combine this with the risk of introducing bugs or missing some obscure bits of functionality, and there is valid reason to object. I find this is particularly common mistake by junior programmers (not saying this applies to you), presumably because they aren't used to reading other people's code. Frustratingly, this is often coupled with an attitude that missing out large chunks of existing user functionality is acceptable if it makes the code a bit simpler. Of course, sometimes rewrites/refactors really are an improvement. Sometimes code really is fragile and confusing, either because of who wrote it or because it has had many small changes tacked on in the easiest places. Or perhaps the last person that understood that code has left the company, so it's OK that you find the code clearer only because you just wrote it! But in any case, it is fair to ask for real justification for a rewrite.
- logicallee 8y ago(Genuine question.) Do you think "small changes" shouldn't be "tacked on in the easiest places"? I'll try to give an example I hope is realistic: One of the heaviest things an application can get is a complete theme system. Suppose you don't have one. It's not a requirement. Adding a theme system when there is none is months of work and might impact basically every line of code that displays anything. So you're not doing it. Now for some exceptional system there is just one case where somewhere you are displaying something under an external widget that doesn't meet its size constraints or whatever - long story short your text is invisible, you want to inverse the font to get it working. You don't have any code like that. Do you think it is OK to add it to the easiest possible place: in this case perhaps you add an optional argument called "need_to_invert_color" (this awkward phrasing tells you it's a hack) to a single function, default it as false, comment it as: //invert the color of the font. This is needed where an external graphing widget with a black background leaks onto our canvas due to not respecting our pixel boundaries, so that our text displays over it. And then where you call comment the same thing, that //currently a bug in the widget code makes the widget leak xyz pixels below its bottom border. As a temporary fix we introduce an argument need_to_invert_color into our display function. As of this writing 3 Jan 2019 we are just using it from here. The correct fix would be for the widget to stop leaking instead, and when that is done white text is unnecessary - and we might not notice. So we start by testing whether the area we will be overlaid over is indeed the wrong color. ETC. In other words a quick hack for a corner case, that doesn't fix the underlying bug (workaround) and even as a hack makes use of something that doesn't exist (a theme system), instead adding and documenting a half-assed thing tacked on.
- quietbritishjim 8y agoOf course this is just my opinion, but in many situations I think tacking one small change on to the easiest place is fine. My comment earlier wasn't meant as a criticism of that practice. It's just I would consider refactoring after this happens multiple times: at some point, all those little tweaks can add up enough that the original design of the "core" code gets lost in the noise. When exactly is the right time can be hard to figure out. Especially because, if enough hacks have accumulated that code really does need a reshuffle, then the job of refactoring is harder, which actually increases the temptation to put it off. But I wouldn't make a big sweeping change just because of one small hack that's a bit ugly.
- z0r 8y agoWriting code is hard, so there's nothing wrong with reading it being hard too. (I'm only half kidding).
- zorked 8y agoFive people understand the system. You refactor it. Now one person understands the system.
- Aeolun 8y agoInitially the system took 8 hours of reading and 2 days of refactoring to grasp. Now it takes 2 hours. I’ll take that.
- chucksmash 8y agoWorking Effectively with Legacy Code recommends just this idea. They call it scratch refactoring. Refactor without a lot of forethought to see how the system works, then just throw away your changes but keep your newfound understanding.
- lugg 8y agoThis is effectively "it." But I do like to branch off the more useful refactors into their own commits.
- pault 8y agoI often see some bit of code, or a dependency in a project that I think is over engineered and too complex and start rewriting; first I get something basic working, then I discover some edge case, then another, and another, and eventually I realize that I have reimplemented the original code. It always makes me feel silly, but at the same time those have been the best learning experiences I've ever had.
- titzer 8y agoTake a moment to consider that someone else in your team might actually understand the code in question, and when you rewrite it (and commit), you destroy their knowledge about how it works. Sure, from your perspective, you made the code better, easier to understand. From their perspective, even if they do the code review themselves, they will not understand it as thoroughly as you do, and they will be burdened with both the knowledge of how it used to be, and how it is now. The net result can be negative, and eventually no one understands any code that they didn't refactor in the last ~6 months, because if it's any older than that, someone's rewritten it. Sometimes you gotta sit down and bite the bullet and read and ask questions.
- blub 8y agoAny company that can rely on an ads cash cow and a large, competent engineering team can probably afford to rewrite most of their software periodically. This won't apply to most companies, hence the conventional wisdom.
- yakshaving_jgt 8y agoLike most things, it depends. There are happy mediums between rewriting absolutely everything and making the fewest possible changes. This is true and viable in organisations of any size.
- paganel 8y ago> team can probably afford to rewrite most of their software periodically I think Google partially does this in order to keep its engineers happy, as you are more happy when you develop something from the ground-up compared to just maintaining something already built. By going down this route those engineers are kept in a “happy state” so there’s less risk of them flying off to other pastures, where they could potentially build the next product that could “kill” Google. Sort of invisible golden hand-cuffs, if you will. I personally find it tremendously wasteful at a societal level but I can see the value of this strategy for Google as a company.
- Aeolun 8y agoWait, huh? How does DRY increase coupling? I mean, I guess the duplicate code/class is now coupled to the two places that use it, but I have a hard time seeing how that is worse than two duplicate instances of the code.
- yakshaving_jgt 8y agoIt increases coupling in exactly the way you just described. Sometimes this is beneficial. Sometimes it isn't. My point is — more often than not — conventional wisdom is that DRY is always preferable, whereas the reality is not that simple.
- BeetleB 8y agoHere is what I've seen (and been guilty of): Two pieces of code in different parts of the codebase are very similar. They have nothing to do with each other - even semantically. But the code is very similar. So someone thinks this is code duplication and creates a function/class/whatever that both pieces of code can use. Repeat all over the place. Then one day, one of those two places needs custom behavior. I can either change that function/class and create complexity (have to now support two use cases). Or I can stop using that function/class in that place and go back to the old solution. Sometimes, this is quite a lot of work as aggressive "DRY" leads to a fair amount of coupling - there could be a few layers of DRY'd code there to untangle. I put "DRY" in quotes because none of this really is DRY. DRY originally was about requirements - not code. No requirement should show up in multiple places in the code base. In this example, even though the code was almost identical in both places, there was little else common. They dealt with different requirements, for completely different reasons. They should never have been refactored to use a common function/class. These days people keep talking about over-use of DRY, but they're really complaining about overabstraction of disparate code - not the DRY in the requirements sense.
- erik_seaberg 8y agoThis gets back to the https://en.wikipedia.org/wiki/Open%E2%80%93closed_principle https://en.wikipedia.org/wiki/Open%E2%80%93closed_principle. I should be able to override just the behavior I want to change, rather than permanently edit the shared implementation eliminating the behavior you want.
- codeflo 8y agoI’m not convinced this works so great for Google. Some rewrites are very noticeable as a user, and things in the UI frequently shift around for no discernible reason. Perhaps worse, they seem unable to get below a certain level of bugginess in products like Google Maps and Gmail. Perhaps because a new round of rewrites always introduces new bugs before all of the old ones ever get fixed. Perhaps their metrics tell them that all of this is fine, I don’t know. But then, you have to realize Google has so much money they don’t really have to spend it very efficiently. For everyone else, this approach to rewrites seems like an extremely expensive way to produce software that’s not even all that great.
- walshemj 8y agoI saw this a BT where on system was rewritten in OWS (Oracle Web Services) used 15 Person Years and around a Million Quid - Not the best use of shareholders money. But some one got to tick some boxes on there promotion track
- heavenlyblue 8y agoWell, but that money was certainly not needed for someone else if it was so readily available.
- walshemj 8y agoLike pay rises, funding more roles to allow more career progression or gasp retuning it to the share holders?
- polskibus 8y agoMany enterprises are rewriting their apps to the cloud stacks for no good reason, at a great capex and later opex.
- stcredzero 8y agoI’m not convinced this works so great for Google. Some rewrites are very noticeable as a user, and things in the UI frequently shift around for no discernible reason. Perhaps worse, they seem unable to get below a certain level of bugginess in products like Google Maps and Gmail. Perhaps because a new round of rewrites always introduces new bugs before all of the old ones ever get fixed. Perhaps their metrics tell them that all of this is fine, I don’t know. But then, you have to realize Google has so much money they don’t really have to spend it very efficiently. The user experience ranges between good to mediocre to bad, depending. Google is simply too rich and powerful to care. So long as the goose keeps laying the golden eggs, they can just keep going along and patting themselves on the back.
- why_only_15 8y agoThis is interesting to me because something like this only works if there are lots of tests and they can be run after every change. If you rewrite code constantly and potentially create new bugs by e.g. not understanding edge cases previous developers put in, then this isn't feasible. With a focus on testing this becomes practicable.
- fergus_google 8y agoIt's common for code at Google to be rewritten without reusing most of the existing tests -- instead, new tests are written for the new code. Yes, this risks not understanding edge cases. But not all of those edge cases are still important.
- why_only_15 8y agoMaybe a compromise would be that if you rewrite old code you individually go through old tests and have to sign off on deprecating them, and say edge case X is no longer important, so you keep most tests while not using ones that don't matter.
- StreamBright 8y agoThis makes sure there is never 1.0 ever. I think this is one of the biggest mistake in software. We just keep rewriting things that are already doing what they supposed to. Like the Gmail UI. It got rewritten 3 times already and every iteration it gets shittier.
- nindalf 8y ago> every iteration it gets shittier I like how you state this like it's an objective fact. I've always been happy with the gmail UI and the latest iteration is great too. Outlook on the other hand...
- Aeolun 8y agoOutlook has always been shit, but at least it has been shit in the same way for the past 15 years.
- StreamBright 8y agoIt is not an objective fact but sampling the gmail users around me gives me an idea. Obviously is it not representative.
- srj 8y agoI think you're imagining much larger rewrites. A system such as Gmail is composed of many smaller parts. If one of those parts was written years prior for a world that has since changed, it may be accruing technical debt as it's continually extended to fit new requirements. An occasional rewrite helps address this type of decay. Without the rewrite you may find yourself 10 years later with a system that's both critical and kludgy, and at that point the rewrite will be a much larger project.
- MichaelMoser123 8y agoI am wondering if the rewriting rule applies to AdWords and search, I mean they probably would get very upset if their cashcow stops working, all of a sudden.
- hknd 8y agoI think we should not take "rewrite" too literally. But both, search and adwords, have been rewritten multiple times since their launch. You can have a look at the papers created by Jeff Dean (1) and Sanjay Ghemawat (2) which mention some of the new concepts/technologies/features used in those products. 1: https://ai.google/research/people/jeff https://ai.google/research/people/jeff 2: https://ai.google/research/people/SanjayGhemawat https://ai.google/research/people/SanjayGhemawat
- MichaelMoser123 8y agoThanks, very interesting articles!
- davidwihl 8y agoThe AdWords API is now being re-written, the biggest change in ten years. It is migrating from XML to gRPC / protobuffers. https://developers.google.com/google-ads/api/docs/start https://developers.google.com/google-ads/api/docs/start Disclosure: I co-author the Python client library and have written a few of the docs on the site listed above.
- MichaelMoser123 8y agoThanks for the link!
- snorkel 8y agoAfter working with banks that still run production code on obscure and obsolete platforms written by people who retired decades ago, I totally endorse this practice ... as long as I’m not doing the rewrite.
- Radim 8y agoThe dance between death by ossification and death by excessive chaos is a delicate one. Be too quick and you're constantly chasing shadows… wait too long and you're immobile.
- sytelus 8y agoIf you have engineers with physiological problem of “not invented here”, you have a very serious issue. I am currently seeing this in real time in one of the projects and I was told almost exact same words as “reason” to recreate what we already have and working beautifully. It was clear to me that some developers are just too lazy to dive in to complex system. They get ticked off by one imperfection here and other over there and immediately run for exit shouting “I could do so much better”. Instead of understanding why things are the way it is, they fantasize about how they can one up original authors and claim their own hero title. They go on to throughly underestimate the time to recreate what has taken years of learning. So they spend next many months sweating out, copy as much code from “old” stuff as they can, dropping important feature here and there, adding new and old bugs - very often arriving at more or less same place they started off. Meanwhile competition has moved on to V2 laughing their way to the bank and customers scratch their heads why you are still stuck in same place for so long. Then our new “owners” gets their promos after massive marketing of how much better everything is now. But to everyone’s surprise they soon leave the project because working on bugs and incremental features has became boring and BTW, the new stuff is just as complex as old stuff. New devs roll in and we start the whole cycle again.
- JustSomeNobody 8y agoI don't think it is laziness. Some developers simply can't ever shake off the need to be the smartest person in the room. They have been told all their lives how smart they are. Now among peers, they're ... well ... average and they haven't been given the tools to understand how to cope with that. Some developers never outgrow that, sadly.
- duncan-donuts 8y agoOne thing I think that is important to point out is the, “working beautifully” part. I agree that rewriting stuff just because is a waste of time and money. There are plenty of solutions out there that may solve a problem and may be working fine, but have now hit their scaling ceiling and people become frustrated. Rewriting stuff is a part of software, but should almost always be done incrementally. I also don’t think it’s laziness. Developer hubris is real and it can get in the way of actual business value. You pretty much nailed everything else. I share the sentiment because I’m currently working on splitting up a monolith into microservices. Most of the “quick wins” are things that didn’t need to be rewritten anyway, and the stuff that actually sucks is difficult to address. It’s tricky to get right.
- austincheney 8y agoI have learned this to be true from work on my open source projects. I tried to explain this concept at my previous employer and my boss looked at me like I was stupid. Stupid makes sense when you justify your existence according to story points. What I have noticed from the travel industry is that backend developers tend to write a ton of original code and rarely refactor anything as though they are scared an improvement is always a regression. Frontend developers tended to not refactor anything either, but then they were scared to write any code at all (note: I am a frontend developer). Part of this fear was justified because their automation was shitty. Often things were properly code reviewed, but changes would only be accepted with the smallest possible footprint. It hurts when the diff is a static text comparison counting every character, which means a bunch of comments explaining things or white space changes looks like a ton of code changes. It also keeps you from removing unnecessary code or reorganizing things. The killer though were frameworks for everything including test automation in various different flavors for the same sorts of things, which is like writing tests for testing of tests. In this case testing became a block to check off that had little or no real value.
- rpod 8y agoHey, I just wanted to say that I really love your gensim library. I've used it for my Master's thesis a couple of years back and it's been a tremendous help. Thanks so much!
- Radim 8y agoThank you! That's so nice to hear :) To get on-topic: Gensim could use such a rewrite as well… The ML world changed, expectations and requirements changed, ecosystem and APIs changed. I changed too.
- na85 8y agoProductivity hack or make-work project?
- dilyevsky 8y agoLots of mental gymnastics to dress up the fact you can’t get promoted without shipping new code.
- titzer 8y agoDisclaimer: I work for Google. I speak only for myself. In my honest opinion, frequent rewrites are by-and-large a disastrously bad idea, for several reasons. If there is one thing I would change about Google, it would be to slow down the frenetic pace of change inside. Rewrites just make the pace of change untenable. And I say this as one who is totally part of the problem: I helped rewrite significant parts of V8, the JS VM in Chrome, particularly the optimizing JIT compiler, TurboFan. (Don't get me wrong--I am not knocking any one specific project, my coworkers, even my leadership, etc). I've been at Google 9 years, and I don't know how barely any of it works anymore. 1. The assumption that requirements and environment around software change so frequently that it must be burned down to the ground and rewritten is a big part of the problem. Why do the requirements of software change? A. Scale. B. Because the software around it changed. Bingo. 2. Rewrites actively destroy institutional expertise. Instead of learning more as time goes on, engineers' knowledge becomes obsolete as old systems are constantly changing and rewritten for unclear benefit. Experts can no longer rely on their knowledge for more than a couple of years. This is extremely bad for critical pieces of infrastructure. In short, no one ever masters anything. This is due not just to incentives but due to change itself. 3. No one ever has time to do an in-depth followup study on whether the rewritten artifact was better than the original. Instead people go on their gut feeling of having rewritten something they often did not write themselves (and did not fully understand) with something new and shiny of their own creation. The justification of the outcome is done, in short, by the people who have a big vested interest in declaring success. (And yes, me too). 4. The idea that the software requirements keep changing around software is promulgated by the exact same people who never spend any time up front simply writing requirements down. Well, no fracking wonder the requirements seem to change somewhere in the middle or years later: they were never anticipated in the first place! We'd do better overall if the industry in general did some good ole requirements engineering. Most people I talk to have never even heard of this. Instead, we never have any time to stop and think about doing things right, but we always find time to rewrite from scratch. 5. Zero incentive to do things right. As a field, as industry, we are actually not very serious about writing good software. Instead, we're just going to trash it after 5 years. So software is constantly bad. But the next rewrite! The drive for rewrites is mostly a swindle in my opinion. The reality is that some software needs to be shot in the head, some needs to be rewritten, but most software needs to be just maintained. That means bugfixes, performance improvements, scalability improvements, and sometimes, yes, refactoring too. But bugfixes and incremental performance improvements don't get anyone promoted. Even more cynically, but very realistically, "old" software that is maintained by experts means a dependency on those experts, and they end up being expensive. Corporations hate when their employees have job security! Rewrites are A.) driven by an influx of young talent who want to make their mark, B.) incentivized by the promotion process and C.) driven by a corporate pressure (everywhere, not just Google) to make sure that programmers and software are commoditized to avoid dependencies, bus factor, and job security.