13 ms·
A Taxonomy of Technical Debt
- ebbv 8y agoCool write up and classification system. A category that affects us is VendorDebt. Things that are inflicted on us by external vendors. Classifying then in a similar manner might help us decide which vendors to dump.
- drawkbox 8y agoAt pretty much every game studio there is an epic internal battle of standard libs vs custom. std::string and [some custom string class] here it is AString is usually the spark. A constant of internal game development is that they think they can always build better strings, lists, dictionaries, collections etc than the standard lib, basically thinking the standard lib is as it was in the 90s and all the work that went into them is bunk. In some cases if you are really pushing memory and not writing custom allocators or using something like boost then yes, but in most cases the technical debt of custom classes written by an ancient from generations ago internally is more technical debt. > One of the best examples of MacGyver debt in the LoL codebase is the use of C++’s std::string vs. our custom AString class. Both are ways to store, modify, and pass around strings of characters. In general, we’ve found that std::string leads to lots of “hidden” memory allocations and performance costs, and makes it easy to write code that does bad things. AString is specifically designed with thoughtful memory management in mind. Our strategy for replacing std::string with AString was to allow both to exist in the codebase and provide conversions between the two (via .c_str() and .Get() respectively). We gave AString a number of ease-of-use improvements that make it easier to work with and encouraged engineers to replace std::string at their leisure as they change code. Thus, we’re slowly phasing std::string out and the “duct tape” interface between the two systems slowly shrinks as we tidy up more of our code. So now there are two string classes, that is technical debt... and one should be consolidated on and the arguments against std::string are sometimes valid but you can also do custom memory allocators or use better standard lib iterations. EA even rewrote the whole standard lib EASTL [1] to adjust for some of these issues i.e. fragmented memory. Some games require it, some it is pure ego in game development teams. Game development teams have the highest ego driven development (EDD) I have ever seen and lots of tricks that take five minutes (but add 2-3 months to testing due to five minute solutions) but are more spaghetti than templates that write templates. The one problem that comes about with your own standard lib or thinking you are better than boost or similar, is that the learning curve on the internal lib replacements add technical debt and start up costs, and the original guy that wrote them is long gone usually. Also, in the end portability suffers as there is invariably 3-4 versions of the internal libs. Developers have to weigh the technical debt of your own custom classes outside standard libs and see if that outweighs the memory issues that may arise. Today most machines are not as affected by memory fragmentation issues and there is more cpu/memory to go around, and where they are you can write custom allocators for std/stl or use something like boost. I do love Riot Games and all game development teams just I have never worked in one or with one that doesn't have the standard lib vs custom battle and wastes lots of time when one isn't standardized on or when not necessary. Some games and game engines require it, where they do you should fully commit one way or the other. Though going custom leads to slowdowns in coding for new devs and invariably there will be multiple versions of those internal libs over time that add up in the debt department. [1] https://github.com/electronicarts/EASTL https://github.com/electronicarts/EASTL
- WorldMaker 8y agoThere's absolutely a "problem domain debt". If you are a game company, your problem domain that gets you paid is your game(s). Time spent rebuilding standard libraries is time not working directly on the game, and maybe time not getting properly "paid". Meanwhile, there are already people paid to work on the standard libraries, and its their job to make those work and continue to improve them. Certainly there are tradeoffs where you may have to know the standard libraries well enough to know their performance characteristics, or how best to mitigate worst case scenarios, but if the people paid to build standard libraries are doing their jobs (that you pay them for when you buy that compiler), it should be less debt work to workaround an existing solution than build one from scratch.
- jrs95 8y agoYou would think the approach taken with this would be to just use the standard libs until they're actually contributing to a bottleneck and then worry about optimizations. Doing it prematurely does seem to be a problem. Though I don't think having custom alternatives to parts of stdlib are bad if you're actually making a meaningful optimization.
- ryandrake 8y agoPremature and/or misplaced optimization. It’s kind of funny that they worried so much about std::string’s performance that they rolled their own, yet have this big honking lua layer thunked on top of everything for game logic. Wow!
- drawkbox 8y agoNot only that, if you want good framerates and 60fps you aren't allocating at runtime, anyone who is doing that at a game dev studio is taken out back and either shot or to work on the cow clicker. Usually standard lib vs custom arguments end up in the weeds like tabs vs spaces at game companies but ultimately it has almost nothing to do with framerate or runtime. Largely it is about that EGO. Why maintain a standard lib instead of improving gameplay and networking? Well some want to be a lord in their feifdom where they are the controller the code and the ring. Riot Games has std::string and AString, but what happens when player two enters the game and you got BString? Then BString invites its friends and you got CString and DString. Now your 'standard' has many standards and is more standardy and like warring lords within internal factions like a Game of Thrones.
- mmsimanga 8y agoIn data warehousing and BI, it's MacGyver and data technical debt all the way down. MacGyver because of all the "urgent" reports whipped up for CEO, duplicate copies of data and the reports done by consultants who barely understand industry. Data dept because of all the bugs and changes passed down as data from source system.
- worldsayshi 8y agoDoes any programming paradigms protect better against data debt? The only way that I can imagine to significantly protect against this would be if there was some way to generate data migrations based on type changes.
- mmsimanga 8y agoI don't think there is any programming paradigm that protects against data dept. It is a question of whether the IT team is willing/able to fix the data, with requisite logging of change. This is always the best solution. It often turns into a political nightmare, business insisting that data needs to be corrected. IT not having the resource, time and the change being too risky.
- LtRandolph 8y agoI definitely agree that keeping data migration in mind at a foundational level can be very helpful. The ability to run scripts/regexes easily against the data can make it easier to reason about the consequences of your data, too.
- hinkley 8y agoI don’t know where it is now, but early versions of Angular encouraged you to isolate all your data debt to the sevice layer. With all the kludges in one place you had a better idea of how bad it was and it was easier to pick a block and insist that it now be handled on the backend.
- paulmd 8y agoI don't think there is any technical obstacle or pattern which can prevent dumbasses from shitting things up, humans are just too creative. As soon as you allow any extensibility, someone is going to start shoving integers in as strings, "Y/N" strings as booleans, etc. It would help a lot if there was a well-formed, unambiguous specification for both sides to hold to. Something like the IETF terminology, in terms of MAY/SHALL, specifying things like "true/false" vs "Y/N", etc. Providing sample responses with decent coverage of the possible options is good as well. Then you at least have the leverage to say "aha but the spec says it should be like this, why are you doing it wrong".
- methodover 8y agoThis is a fantastic article. Contagion is a really great term. I've seen my poor abstractions be replicated by others on my team, to my horror -- "don't they see why I did that in this particular case, and not in this other case?" Of course, that's entirely, 100% my fault. I picked a poor abstraction, I put it in the code, I didn't document it well enough, and of COURSE other programmers are going to look to it when solving similar problems. They should! That said... Sometimes I spend a bunch of time finding the right abstraction for a feature that we end up not expanding. And then it feels bad that I spent all this extra time coming up with the "right" solution, instead of just hacking out something that works. Hmm...
- zachsnow 8y agoI have found the closer I am to the product and the clients that will be affected, and the more thoroughly I understand the usecase from the client’s perspective, the better I am at understanding how much effort to spend on “getting it right” in this way. Still wrong sometimes though!
- worldsayshi 8y agoInteresting how you point to a slightly different kind of contagion in replicating code patterns. While the article seems to discuss the kind that is inevitably forced on whoever depend on the code.
- theptip 8y agoI found contagion to be a great clarifying concept too; it's something that I've been looking at in my codebase as the team expands. My gut feel is that it's not necessarily about what you write in the first place, but what you refactor -- sometimes you can get away with a gradual replacement strategy (like std::string => AString from the article), but if the original pattern is contagious and bad, then you might have to take a more aggressive one-shot refactoring approach. I've definitely seen this where a localized refactor is made to try to find a better way of doing something, we decide that we like the new way, and then don't find the time to replace the rest of the usages, resulting in a confusing state of affairs where you need to know which is the "blessed"/"correct" way of doing things. I think that "contagion" is a good lens to use when assessing what the refactoring strategy should be for a given change to the codebase.
- jtchang 8y agoI love this article. Quickly breaks down the types of debt.
- jedanbik 8y agoReminds me of risk analysis: Impact times Probability equals Risk. Contagion seems like a probability factor. Impact is the cost of leaving things unchanged. Fix cost is the cost of fixing the problem. Risk management in this context then means comparing Impact cost to Fix cost in terms of impact for the business.
- stouset 8y agoThe one difference is that contagion is multiplicative over time (potentially logarithmically, linearly, or exponentially—probably a reasonable definition for 1/5, 3/5, and 5/5 respectively).
- billysielu 8y agoI find it's always worth asking "will this get better over time, or worse" for everything, ever. Folks just fail to see past the next few months, having at least one person in the room asking this question makes them at least ignore it intentionally instead of complacency.
- deleted 8y ago[deleted]
- jimmaswell 8y agoSomewhat aside, but the brain having to "flip" visual information because it's "upside down" seems suspect to me. Turn it sideways while maintaining all the connections it has to the rest of the body, and what changes? Is it getting visual information sideways that it has to rotate now? Probably not.
- ninkendo 8y agoMoreover the idea that the collection of neurons that your retina connects to has any concept of "orientation" is nonsense to begin with IMO. It's not that "there's an upside-down image that your brain has to fix", it's just that your brain interprets signals from your retina as a picture in your mind, full stop. Rods/cones in the top of your retina connect to your brain through neurons, so do the ones at the bottom. But to say that "this 'top' retinal cone should really connect to a 'top' neuron in your brain", doesn't even make sense to me. Since when do the locations of the neurons interpreting the input even matter? It would be the same with hearing too... you have a left and right ear, but if for some reason those were swapped and your left fed things to the right half of your brain and vice-versa, your brain wouldn't be "flipping it back", because how could the absolute location of the neurons interpreting the sounds even matter?
- evanwise 8y agoThis is the right way to look at it. In fact, your brain is plastic enough that if you wear glasses that flip your vision upside down for several days it will eventually relearn the mapping of retinal cells to neurons so that you see things normally while wearing them. This was studied in the 1890s by a guy called George Stratton.
- jpfed 8y ago"Since when do the locations of the neurons interpreting the input even matter?" Incidentally, these neurons theoretically could go anywhere (as long as they're connected correctly), but in practice they end up arranged retinotopically (https://en.wikipedia.org/wiki/Retinotopy https://en.wikipedia.org/wiki/Retinotopy).
- 8y ago
- mitko 8y agoGreat article, loved how the examples were presented. In my time as an engineer, I've found that thinking of tech debt as financial debt also helps. There is the initial convenience (borrowed money) of using the debt-ed approach. Then there is fix cost as Bill Clark name it, i.e. how much to pay back the debt if it were money. The impact is akin to the amortization schedule, i.e. what is the cost every time. For normal money, amortization schedule is over time, but for tech debt it is over usage. The amortization schedule of tech debt is discounted over time, as with money, _now_ is more important that _later_. Contagion is a great concept, and I think it is a better name than interest rate, as the debt will spread through the system, and not just linearly with time. Tech debt is also multi-dimensional and not fungible like money, which makes it a harder thing to reason about. But the good news is, in my opinion, that sometimes it is perfectly fine to default on some tech debt, and never pay it back, delete the code. Then taking that tech debt was a win, if the convenience was more than the amortized payments.
- baddox 8y agoI think the main difference is that technical debt is not fungible, i.e. you can’t necessarily easily choose to pay off the highest-interest technical debts first like you would for your personal financial debt.
- oculusthrift 8y agoput another way: you can have one item that is 5 days of work but really critical and another that’s 2 days of work but way less critical. If you have 2 days to work on tech debt, you basically are forced to do the 2 day one. especially since you are evaluated on what you finish, not how much you worked towards some long goal.
- baud147258 8y agoAs financial analogy, I've seen a piece (linked on HN a few years ago) comparing technical debt to unhedged options, meaning you can get a benefit and you might or might not get bitten by it.
- humanrebar 8y agoI might have missed it, but missing from the taxonomy: "Pay In Full" Debt. In this debt, you pay the entire cost until the last use of it is cleaned up. This kind of debt is especially insidious because there is no incremental benefit to cleaning it up.
- piinbinary 8y agoI'd be curious to hear an example of that (I don't believe I have personally seen one that fits that pattern in the wild yet).
- cbanek 8y agoI'm not sure if it's "technical debt" but backward compatibility between versions can feel like this. Once you decide you stop supporting a previous version, you can rip out all that code (or strange compatibility code paths). Until then you're stuck with the whole thing.
- oculusthrift 8y agoremoving a library? you can remove 100 usages of it but not until every single one is gone can you remove it.
- TremendousJudge 8y agoporting to a backwards incompatible a language version? you can't use most of Python 3's new features while some part of your codebase is in Python 2
- logophobia 8y agoI used to work at a large software company. They had one giant monolith application, about a hundred different modules that did things from financial registration, displaying media. 30 years old, ported from clipper to delphi to .net. Lots of technical dept, but one issue fits the type of "pay-in-full". They only had a maximum of 3 gigs of ram to work with because the monolith was a 32 bits application. That was ok for most modules, but some did some rather complicated stuff that required more memory occasionally. It caused infrequent out-of-memory exceptions. The cause was one 32-bits library that was very hard to replace. It was used everywhere, and stored reports in a proprietary format, the company that made that library went out of business at some point, no source was available. The company prided itself on backwards compatibility, so it couldn't just dump the library without porting all the reports over. As far as I'm aware, it's still a 32 bits application.
- lifeisstillgood 8y agoThere are writers who just ooze technical depth of understanding - i thinks it's something to do with trying to explain something at a laypersons level, but leaving many assumptions just there for the reader to follow. It's almost the opposite of baffling with bullshit. Good read and a really useful concept
- kashyapc 8y agoThis reminds me of the following, from the book Team Geek[1], chapter "Offensive" Versus "Defensive" Work: [...] After this bad experience, Ben began to categorize all work as either “offensive” or “defensive.” Offensive work is typically effort toward new user-visible features—shiny things that are easy to show outsiders and get them excited about, or things that noticeably advance the sexiness of a product (e.g., improved UI, speed, or interoperability). Defensive work is effort aimed at the long-term health of a product (e.g., code refactoring, feature rewrites, schema changes, data migra- tion, or improved emergency monitoring). Defensive activities make the product more maintainable, stable, and reliable. And yet, despite the fact that they’re absolutely critical, you get no political credit for doing them. If you spend all your time on them, people perceive your product as holding still. And to make wordplay on an old maxim: “Perception is nine-tenths of the law.” We now have a handy rule we live by: a team should never spend more than one-third to one-half of its time and energy on defensive work, no matter how much technical debt there is. Any more time spent is a recipe for political suicide. [1] http://shop.oreilly.com/product/0636920018025.do http://shop.oreilly.com/product/0636920018025.do
- LtRandolph 8y agoOooh, I like that a lot. Thanks!
- hinkley 8y agoThe XP guys had it right. Amortize all defensive work across EVERY piece of offensive work. In the tech debt parlance most people are paying interest only payments instead of paying against the principle. Every check you write should do both (extra payments are good but they aren’t good enough).
- hywel 8y ago"I’ve rarely encountered discussions of contagion." This surprised me: contagion is a good metaphor because it is a compounding measure of the growth of the problem. Just like an interest rate (a compounding measure of the growth of debt). Most senior developers I've met have considered the interest rate of the debt, which seems like it has been renamed here as contagion. Maybe I've been lucky to just know smart people! From the point of view of explaining these concepts, I'd suggest keeping the metaphors consistent. Tech debt should have an amount owed and an interest rate, tech infection (?) should have a potency and a contagion level.
- matte_black 8y agoI would love the idea of a technical credit score. For example, if you’re the kind of dev that racks up technical debt and never pays it down, you should have a shitty technical credit score, and be considered a poor hire. Whereas someone with great credit, would be a great asset to bring onto the team.
- mic47 8y agoTracking "credit" score sounds like good idea, but I would not go as far and assuming that persons with bad credit scores are poor hires. Maybe person who creates tech debt is really great at prototyping, fixing urgent issues with unconventional methods (aka MacGyver) or do other tasks you find boring. While credit score of this person will be low, such people are also great assets in the team. In general, this metric could be useful as tracking number of pull requests, lines of code, and so on: to spot anomalies and investigate: maybe that person is suddenly blocked by something, overwhelmed and need help, or just works differently, or on different tasks and the anomalous metric is ok.
- brightball 8y agoA metric like that would also discourage people from actually documenting the debt. “When a metric becomes a target it ceases to be a good metric.”
- rickbad68 8y agoIn my experience often the term 'technical debt' will be hijacked by product-oriented folks resulting in feature debt being presented as tech debt.
- scarface74 8y agoI've been binge listening to Software Engineering Radio for the past few months. I am currently listening to an episode where they are talking about technical debt. He has the opinion that clean code is not as important as shipping code - ship the code first and then refactor as needed after you get customers. http://www.se-radio.net/2015/04/episode-224-sven-johann-and-eberhard-wolff-on-technical-debt/ http://www.se-radio.net/2015/04/episode-224-sven-johann-and-...
- nkristoffersen 8y agoDefinitely. I build a lot of MVPs. So I focus on "make it work, then make it work well". Shipped code is so much more valuable than unshipped code :-)
- jeffdavis 8y agoWhat about "fear"? The most pernicious thing about technical debt, in my opinion, is that it creates fear in the sense of "I don't want to touch that module". Even if you try to be objective and use hard facts to overcome the fear, it doesn't matter, because fear destroys creativity, so you've already lost.
- kraftman 8y agoYour tests should reduce that.
- arca_vorago 8y agoThis seems far too focused on dev tech debt, which has a very narrow scope. I like the article, so I'm not knocking it, just offering a little perspective. As a senior sysadmin in the past my primary issues have been technical debt across the entire board, number one being too few hires for too much workload due to cheap or nearsighted execs, but I would definitely agree that contagion is a great term for how techdebt grows faster the longer it's left alone. It's worth remembering the CTO and senior sysadmin and a few others are dealing with all the tech debt of the entire company and IT department of which dev is only a subset (of course this depends on the company, but on HN sometimes I see convos like this where it feels devs are just talking at each other and not receiving much outside feedback.)
- LtRandolph 8y agoI'm not surprised at all to hear that I have blind spots outside of "dev". I've been working on shipping games for a decade, so I'm very fixated on the types of stuff I run into day-to-day in that dev process.
- arca_vorago 8y agoNothing wrong with that at all. Has that all been on LoL? You guys did a lot right with it from the start. Unfortunately I burnt myself out on mobas due to HoN.
- eadmund 8y agoIt's a great article, but I do have one quibble. > A hilariously stupid piece of real world foundational debt is the measurement system referred to as United States Customary Units. Having grown up in the US, my brain is filled with useless conversions, like that 5,280 feet are in a mile, and 2 pints are in a quart, while 4 quarts are in a gallon. The US government has considered switching to metric multiple times, but we remain one of seven countries that haven’t adopted Système International as the official measurement system. This debt is baked into road signs, recipes, elementary schools, and human minds. A not-so-hilariously stupid mistake is to think that the traditional measurement system is stupid. His picture illustrates one of its virtues: the entire liquid-measurement system is based on doubling & halving, which are easy to perform with liquids. The French Revolutionary system, OTOH, requires multiplying & dividing by 10, which is easy to do on paper or with graduated containers, but extremely difficult to do with concrete quantities (proof: with one full litre container and two empty containers, none graduates, attempt to divide the litre into decilitres). The real foundational debt is that we use a base-10 system for counting, due to the number of fingers & thumbs on our hands, rather than something better-suited to the task. If we fixed that problem, then suddenly all sorts of numeric troubles would vanish. There's actually a lot to be said about the Babylonian base-60 system, to be honest.
- baud147258 8y agoWhat's better about a base-60 system compared to a base-10 system?
- plopz 8y agoProbably the same benefits as a base-12 system compared to base-10. More divisible factors.
- TeMPOraL 8y agoThat's an... interesting point I haven't seen brought up before. Makes me appreciate the "traditional" system more. Still, I guess we aren't going to drop base-10 any time soon, so I believe the US should just accept the "traditional" measurement system as something that used to be very practical, but no longer is due to progress of technology, and switch to SI.
- debt 8y ago“We gave AString a number of ease-of-use improvements that make it easier to work with and encouraged engineers to replace std::string” Are you absolutely sure this itself won’t become Foundational technical debt? You seem overly confident, given the metrics, that replacing std::string is a good decision.
- LtRandolph 8y agoWe certainly can't know for certain. But we've had a significant, measurable reduction in CPU cost due to "hidden" memory allocations from things like passing a char* into a function that takes a std::string and stuff like that. (I may be being mildly inaccurate, as I wasn't the guy doing the perf captures etc. I just talked to him about it). I'm particularly impressed by AStackString, which is a subclass that has initial memory allocated on the stack, but automatically converts to dynamic allocation if you exceed that space. So we get quick stack allocation by default, but it will safely handle when it needs to expand. Most of the quality of life stuff is around having in-built support for printf style formatting, string searching (including case-insensitive).
- carapace 8y ago(MacGyver's name is Angus!?)
- monkeydust 8y agoI am a senior product manager for a large financial technology company. Over the years I have learnt to become comfortable with allowing my engineering teams to refactor code whilst delivering new functionality. This has been a process and largely one of trust between me and the engineering leads. It has also helped that I have seen payback from the investment made from reducing down the debt in terms of us delivering new functionality quicker and less error prone code. Although, this payback can take a while to see (6months + which is a long time for a product person operating in a competitive space!) Most of my managers don't get this or if they do they are too blinded by immediate kpi's from further above they can't justify it so in most cases I just tell the engineering guys to add a spread to their estimates to cover the paydown of the debt. Over they years this has definitely helped me build tighter relationships with engineers which as any product manager knows can have huge benefits.