12 ms·
The ‘flawed five’ engineering productivity metrics
- dtagames 4y agoThe most useful metric in my view, and one I learned at IBM, is fixes applied over shipped lines of code. Multiple fixes over the same lines of code is exponentially bad. IBM started measuring code defects vs working code in this way because productivity studies they did showed that fixes took much more time per LOC than new code and had other costs (customer sat, doc changes, reputation) besides.
- wahnfrieden 4y agoThis is a variant of DORA’s change failure rate
- renewiltord 4y agoBut ultimately, the outcome is that IBM is a dinosaur corp and no one looks to them for technical leadership of any sort. They don't deliver and they're not well known for writing particularly good code. So that calls the value of the metric into question.
- prepend 4y agoThis seems easily gamed by just producing lots of lines of code. Adding inline documentation would improve this metric. This does explain some absolute dogshit products IBM made as maybe they were optimizing for this metric by having 1000 lines when one would do. I’m bitter from having to decompile and debug websphere in the 90s and 00s. I think this runs into a problem is that programmers are good at minmaxing. So any rote metrics will end up being gamed pretty quickly.
- dtagames 4y agoEverything was peer reviewed, so no extra code. Also, every single line had a comment, so no padding with comments. On the "why they failed" aspect, 100 years is a pretty good run. There are lots of reasons IBM is less relevant today, but buggy software isn't one of them.
- prepend 4y agoI’m sure it’s buggy software didn’t help and I think it contributed to IBM’s reputation for poor software.
- xyzzy4747 4y agoThe best metric is your own intuition about how productive people are being and their output, subjectively thinking about the quantity, quality, and impact. It becomes pretty obvious who is contributing a lot and who isn’t. You don’t need to track metrics.
- Traubenfuchs 4y agoThis also doesn‘t work and just gives social, extroverted and eloquent people a huge advantage.
- xyzzy4747 4y agoNot really. If social and extroverted people provide more value to the company, they should get better performance reviews and paid more.
- anotherhue 4y agoIt's probably easier to measure anti productivity than software productivity. Excessive or poorly timed meetings, scope change, poor WFH distraction management, excessive support load, etc. I'd like to think that we can assume people will be productive if we set them up for it.
- ptudan 4y agoAmazon promos and firings are based a lot around the amount of lines of code you write, the number of code reviews you do (and the percentage of the time you review when asked), the number of merge requests you have, and the number of iterations per review. If you average more than 2 iterations per MR, you're on the chopping block as it means you're "sloppy". Its ridiculously dumb. I've heard those numbers matter less as you gain tenure and seniority. But for the new grads and lower level engineers, all the savvy ones were gaming these metrics. It sucked to work in that environment and was a big reason why I left.
- twblalock 4y ago> Its ridiculously dumb. I've heard those numbers matter less as you gain tenure and seniority. But for the new grads and lower level engineers, all the savvy ones were gaming these metrics. The company care more about weeding out the really bad junior engineers than it does about rewarding the good ones. Honestly I don't have a big issue with that because I've seen what happens when bad people stick around. It's never fun for the good people, though.
- ptudan 4y agoWeeding out bad engineers is necessary. But these metrics are usually used to justify firings on the face than actually using them as an evaluator. Then the directors can pat themselves on the back that they fired someone justly.
- hammock 4y agoMy high school summer jobs was working the phone at an inbound call center, where there is a constant queue of customers calling and you answer each call in turn. The software used would spit out individual metrics like "number of calls taken," and our supervisor used to look at that metric to make sure we were working and not slacking off at our desks, not on the phone. People figured it out and would just pick up and hang up the phone ten times in a row to game the metric. In response, the supervisor started looking at "average call time" in addition to the number of calls, to make sure this wasn't happening. So what people did instead would be pick up ONE call, and leave the line open for an hour, long after the customer had already hung up, in order to game that metric. Seems like something similar could be done with these Amazon metrics.
- jph 4y agoIMHO teams works best when they choose their own key performance indicators, and match these up with the real-world success of the team's users, customers, and stakeholders. These kinds of metrics can involve people (e.g. add a feature to increase customer satisfaction by X points), performance (e.g. optimize a path to increases throughput by Y%), processes (e.g. fix a bug so security continues to match commitment Z), etc.
- dijit 4y agoThere's some for SRE's and Sysadmins too: * Cost reduction; usually by a very arbitrary amount, despite you having no control over what's needed * Uptimes; last job told me that I had to get 99.998% uptime, Googles global load balancer is only 99.9%
- MichaelBurge 4y agoYou might as well just promise 100% uptime. If you don't meet it, most SLAs you're only liable for a couple bucks in service credits anyways.
- dijit 4y agoI'm mostly talking about a metric by which my teams performance would be judged. Externally to customers we had no promises of availability.
- yodon 4y agoMy favorite quantitative metrics for engineering teams: - Avg time from code review requested to code review picked up - Avg time to complete code review - Avg time from eng done to first customer using it - Avg time from eng done to full production release - Fraction of tasks started that never reach a customer These are loosely based on the Japanese concept of Muda (waste), as personified in the physical logistics world via the acronym TIM WOOD (or TIM WOODS)[0] and are similar but not identical to the DORA metrics. The time to complete a code review is there (for example) not to focus on the amount of time actually spent performing the code review but to focus on all the waiting around the actual code review, which is typically much much longer than the time spent doing the review itself. It's not uncommon to see organizations where an engineer will submit code for review and then have to wait a day or more for someone to pick up their request and then another day or more for that other person to get around to reviewing it. If there are comments on the commit that need to be responded to, you can see additional delays. These "minor inefficiencies" can have huge impacts on the poor dev who is trying to get their code merged, and cumulatively they result in significant increases in feature latency, the total calendar time required to ship a feature. [0] https://www.shmula.com/28695-2/28695/ https://www.shmula.com/28695-2/28695/
- kqr 4y agoThese are great! How are you using the last one? I feel like fraction of tasks started that reach customer could easily become misleading: you want to quickly abandon tasks once you realise they're no longer viable. In fact, the development effort should be partially about finding reasons to stop working on the thing, so you can toss it out as soon as possible, instead of waiting for the customer to turn out not to use it. I would even say that canceling many started tasks is directly correlated with a quick cycle time, by Little's law.
- yodon 4y ago> Why are you using the last one? The last one is definitely in a different bucket for me than the first four. For starters, all of these are there primarily to encourage conversation. That said, the first four can make a lot more sense to try to graph and track and optimize. The last one tends to be more purely about driving a conversation. At the top level, if you're genuinely doing a ton of learning along the way to shipping the right feature to the customer, then arguably that value is ultimately reaching the customer. If on the other hand you can't decide what the goal is and you keep changing your mind (as is often the case), then you tend to end up with a lot of dev investment made in things that simply never ship. It's also worth differentiating technical "spikes" from feature "experiments." In my vocabulary, spikes are things where you're internally assessing a question like "could we do this" and experiments are things where you're externally assessing "do customers want this/does this have the impact we want." If you have a lot of experiments that don't reach customers, you're burning a lot of dev time on things that aren't actually experiments (because by this definition experiments need to reach the customer surface to deliver data). That's a signal you should probably be looking at. Spikes generally only reach the customer indirectly (through an eventual shipping feature), but if you have a lot of spikes that don't ever reach the customer in any way that's also a signal you should probably be looking at.
- shakezula 4y agoFor fuck’s sake, can we just stop trying to measure developer productivity like we’re an assembly line?
- ok123456 4y agoBut then how will non-technical managers justify their salary?
- ethanwillis 4y agoI know the question is probably sarcastic. But being serious: Simply by being good at their jobs of helping the people they manage and having those people want them to be part of the team. And then on the flipside as well where higher ups trust them to help be a good translation layer that enables teams to meet organizational objectives.
- hinkley 4y agoIf you have to chose between two managers based on metrics, then you're already fucked one way or another. Either because you can't actually afford to lose either of them, or their both so awful that just asking people doesn't get you a good answer. I've said it before and I'll say it again: finding ways to characterize people as bad at their jobs is about keeping salaries down, whether by accident or on purpose. Because those are metrics your boss's skip level manager looks at. And they 'work' until they don't, by which point the manager can move up or out and get a reset on the numbers being used against them.
- brazzy 4y agoSo how do we measure it?
- shakezula 4y agoWhy do we need to? No, seriously, why? I have yet to see any meaningful increase in a team’s productivity after they start tracking “developer productivity”. Each time it results in a blow to developer morale and a pretty dashboard that management uses to retroactively justify their decisions.
- kodah 4y agoThe problem with all of these metrics is that they assume that there's an engineering team, or someone on an engineering team, somewhere that's not doing shit. Whether the business believes they're not doing shit because they refactor more than they write features, because they release twice a year instead of every week, or because they have less frequent merges. The entire emphasis of measurement is squarely a technical one which comes from the (management) belief that some engineer somewhere isn't contributing and that's bringing the team down. I've rarely encountered these kinds of engineers, much less teams, and designing an incredibly painful system that senselessly costs people their jobs and livelihood all in the hapless pursuit of identifying them seems fraught. Where these metrics would be useful is if the business actually looked at itself first when teams or team members underperform. In all honesty, I've never met a manager with this kind of mindset; they're usually captured by the belief of the above in some way. Businesses do have an old way of determining whether a team is meeting its goals: KPIs. If the team is responsible for a succinct domain, problem, or stack then these KPIs are easy to draw and measure against because they reflect business outcomes rather than trying to normalize for how everyone on a team contributes.
- prepend 4y agoI’ve worked with teams and individuals that do nothing. Like literally nothing, they do shit. I once had to wrap up a product release for contract close out or something and it involved getting all the code and commits from developers who were rolling off. It amazed me how many had zero code they had written in months. I had one developer that had never committed anything in the three months he was there. Obviously this is a management problem and it’s not a single individuals fault. But the manager had like 90 contractors reporting to them and didn’t care that people had zero lines of code written. The developer’s job was to code. I wanted to mention that as there are developers and designers who don’t code but are productive in other ways.
- kodah 4y agoI believe you, I'm sure they exist, my argument is that they're not common enough for this level of pain. Personally, the way that org is sounds like it was by design.
- jrockway 4y ago> Then there’s the naming of it. Calling a metric ‘Impact’ sends a strong signal about how it should be used, particularly by managers. And this makes it very easy to misuse. Impact is pretty clear to me. You can find the site of an impact by looking for the smoke coming out of the crater.
- teeray 4y agoI’ll add another one: code coverage. Coverage is, at best, a proxy metric for how easy it is to test your code. If you make it easy to write tests, coverage generally takes care of itself.
- koliber 4y agoDecision making purely based on such metrics is wrong. It's management by numbers, and similarly like coloring by numbers, while relatively easy, will not produce great results. At the same time, metrics do have a place. Even flawed metrics, like the ones this article describes can provide value. When used together with qualitative evaluation and thoughtful analysis, it provides a more complete picture of what is going on in a team. Metrics such as these provide an addition perspective on a team. A manager knows what their team should be doing. A good manager should have an intuitive feel of what is going on. A manager should have a good qualitative idea of how their team is doing. If the metrics do not align with the other perspectives, something may be off. If a manager believes a person should be coding, the person is not bringing up any challenges, is reporting progress, and they produced 3 small commits over the past month, it is time for a conversation to find out more about what is going on.
- kqr 4y agoYup. Performance measurement is like planning: the outcome (metrics, a plan) is useless. The process you take to getting there (discovering, questioning, measuring, imagining, simulating) is everything.
- hinkley 4y agoMetrics are for asking questions, not answering them.
- koliber 4y agoI like this a lot. Did you get that from somewhere, or did you come up with it on your own. It's captures the spirit of my metrics philosophy really well.
- Normal_gaussian 4y agoWhen leading these are the metrics you should care about: Oldest MR - this should always be less than 2 weeks. This should normally be less than 1 week, but its not worth caring about at less than 2. Unfinished sprints - sprints should finish with enough time left over to cope for an incident in the week. The extra time should be used for planning and continuous professional development (CPD). When someone is trapped in overflowing sprints it means they are deprioritising and undercompleting work that will come back to bite you. Track other metrics for at most 3 months each, ideally only a month or a single week spot check. This prevents gaming and obsession whilst letting you reason about more nuanced behaviours.
- itsdrewmiller 4y agoYour comment is getting some downvotes despite having interesting metrics that aren't super commonly discussed. I bet if you presented it as "Here are some metrics I have found useful" you would get a much more positive reaction.
- Copenjin 4y agoI can't say that I've never used the first three to spot people with absolutely no useful output, sadly they are pretty good metrics for that.
- imwillofficial 4y agoI noticed no solution was offered. Some metric ton s usually better than no metric
- btrettel 4y agoIn the past, I was a patent examiner at the USPTO. Patent examiners have their own problematic performance metrics. In Oct. 2020, the metrics went through a big change that made them a lot more complex. I suspect that part of the motivation was to make the system more opaque so that it'd be harder to game. But in practice I think it added just as many if not more ways to game the system. I quickly figured out that under the new system, you could increase the amount of time you get for a particular patent application through a particular reclassification procedure called a C* (pronounced C-star) challenge. I'm surely not the only one who figured that out. The reason the C* challenge exists is to reclassify a patent so that it can be transferred to a more qualified examiner. But if it's not transferred then the amount of time you get can be changed. That's not necessarily nefarious as many applications have the wrong classification and would give you a lot less time than if they had the right classification. But examiners aren't incentivized to switch an application to the right classification. They're incentivized to change the classification so that the application gets transferred or change the classification so that they get more time. In the latter case I'd intentionally avoid adding (or even delete) any classifications that would reduce the amount of time I got. I don't suspect the long-term dynamics of this system are what USPTO management intends.
- foolfoolz 4y agovelocity points is not useless. it can’t be used in isolation but points delivered by individuals is a great starting place to identify outliers in your org. generally if someone is delivering far higher or far fewer points they are making an outsized impact to the team (either positive or negative). it’s not perfect, you must take context with it, but with averages and on long time scales it’s quite reliable
- prepend 4y agoI think velocity points are useful within a team over time. They are locally useful. But they are stupid to measure across teams or to compare teams or productivity. Velocity points are just an estimating tool, not a measure of value. It’s useful to know that a team usually produces 10 points per sprint but this sprint is 5 or 20. It just lets you know if your team is producing “normal” or not. It’s useless to try to calculate that out of 20 teams the average velocity points are 10 per sprint.
- foolfoolz 4y agoyes. sprint points do not translate outside team scope
- Traubenfuchs 4y agoThey are COMPLETELY useless and are actively gamed by clever devs to reduce workload and reduce output expectations. I always nudge fellow devs to overestimate the tickets I will be working on by overstating the complexity and risks, sometimes I prime them with higher numbers, etc.
- madcaptenor 4y agoMy kid is figuring out pooping in the potty. We have a chart where we make a check mark when she does it. She likes check marks, especially when she can make them herself. Over the past couple weeks her poops have gotten smaller and more frequent.
- cwilkes 4y agoSounds like she’s going up for promo!
- madcaptenor 4y agoShe already got promoted to big sister and we are not having another one any time soon.
- rightbyte 4y agoYou should make bigger marks for bigger deliveries to not give the wrong incentives.
- prepend 4y agoDoes anyone actually use these as continuous variables and evaluate them. I’ve worked for 10 orgs for almost 30 years and while these existed, I’ve never even heard someone propose to use them to measure productivity. #commits are useful as a binary metric that a developer is alive, but trying to say one is more productive than another because they had more commits is pure madness that permeates an org so that I would detect it during an interview and avoid.
- skeeter2020 4y agoWe track all of these and publish them on a continual basis (dashboards), with the exception that we capture deployment frequency not commit frequency. They're directly if weakly correlated. but deployments is closer to what you care about. I don't think anyone is saying you should only look at the metrics and not the qualitative factors (what's in all those frequent commits?) but they definitely help drive conversations and decisions. The alternative (pure qualitative & gut feeling) is much harder to get consistent across an entire engineering department.
- dandare 4y agoOne thing that baffles me in corporate IT is not just the snake pace of development but rather the fact that nobody seems to be bothered by the snake pace. There is zero effort to measure or speed things up. The only important thing is to be nice to everyone, any mention of productivity is considered hostile behaviour. (For reference, I am talking about cases where a team of 5 devs takes 2-3 months to deliver a feature that would take a single independent developer maybe 2-3 days.)
- Traubenfuchs 4y agoA slow pace makes it possible to slack off more in peace. If you plan to take a week to do feature A and you finish it in a day, you have 4 free days. If you plan for 1 day and it takes 2 because it was harder than expected, plans get messed up, you need to work faster on the next feature and look bad. I always encourage fellow engineers to vastly overestimate tickets. That‘s also important to set a comfortable pace with the business people who have zero clue how hard our work really is and prevent them from making us work hard.
- a1369209993 4y ago> the business people who have zero clue how hard our work really is Note that this goes both ways; the'll both wildly overestimate and wildly underestimate how hard various tasks are, often in the same conversation. See eg https://www.explainxkcd.com/wiki/index.php/1425 https://www.explainxkcd.com/wiki/index.php/1425.
- rightbyte 4y ago> snake pace You mean snail or like moving in an "S"?
- jqcoffey 4y agoI’m surprised that folks are still considering metrics like LoC and commit frequency to measure developer productivity, even more so due to the (anecdotal, from my XP of 25 years in industry) fact that as developers gain in seniority they are typically spending more time with people than with code. IMHO, developer productivity is best judged by the humans they work with.
- skeeter2020 4y agoas long as you're looking at the content of frequent commits I think this is valuable as it encourages smaller task sizes. If you use a PR/MR approach it also corelates with other important metrics like WIP and how long the coordination work takes. LoC is not something I was aware people are still tracking.
- vannevar 4y agoI think the real message here is not that metrics are bad, but that they are misused. Imagine if every time you went to the doctor with a fever and they took your temperature, the doctor prescribed an ice bath to bring your temperature down. You wouldn't conclude that thermometers are evil, you'd switch doctors. Same goes for most of the metrics here. Metrics are useful to navigate BY, not to navigate TO. If you have skilled and experienced managers, you can get a lot of value out of all of the metrics listed in the article.
- jameshart 4y agoAll of these metrics are like trying to measure progress on a building project based on the volume of noise produced. ‘I don’t hear hammering! There should be more hammering!’ ‘Lines of code’ is a useful metric for ‘likely ongoing maintenance cost’. ‘Impact’ is a good proxy measurement for ‘likelihood the change introduced a bug’. If you encourage teams to increase those numbers you will get what you deserve.
- bawolff 4y agoDid anyone i the last 2 decades ever think otherwise? These aren't just flawed but some of the most infamously flawed metrics.
- ftio 4y agoI worked as a PM on internal developer productivity at Google for a few years. As I've said in previous comments, compared to my former colleagues, I'm an infant in this area, so take this with a heaping of salt. (Opinions my own.) I do not believe in the possibility of a "General Theory of Productivity," and management-by-numbers-alone is actively harmful, but I do believe in the possibility of measuring productivity in a useful way. Even "bad" metrics like commits per engineer per week can be useful at the right granularity, e.g., to do high-level velocity forecasting over a large, representative group of engineers during different times of year. If you're wondering: different metrics are suited for different use cases, but as a baseline, I think the DORA metrics[1] are a reasonable starting point. 1. https://cloud.google.com/blog/products/devops-sre/using-the-four-keys-to-measure-your-devops-performance https://cloud.google.com/blog/products/devops-sre/using-the-...
- owlbite 4y agoSame old story - as soon as you start using a metric to incentivize people, they optimize to the metric. If the metric is not well aligned with what you actually wanted (and I mean optimizing it is what you want, not currently correlated before you incentivized it), you are not going to be happy. It always amazes me how otherwise very smart people don't think through the consequences of "paid by the X".
- gary_0 4y ago"When a measure becomes a target, it ceases to be a good measure." - Goodhart's Law. See also, the Cobra Effect: https://en.wikipedia.org/wiki/Perverse_incentive https://en.wikipedia.org/wiki/Perverse_incentive. If you set up an incentive system, people will (perhaps not even consciously) start trying to game that system. Maybe there's some way to use machine learning to turn performance metrics into a black box that considers every possible data point? (Now there's a terrifying idea.)
- hinkley 4y ago> So not only is this metric inaccurate, but it incentivizes programming practices that are a counter to building good software. There is a critical error in this statement that some would label as 'subtle' but it's about as subtle as a rusty axe to the forehead. Judging people by lines of code doesn't 'incentivize [] practices that are counter to building good software'. It punishes you for writing good software. While we do need to be aware of the negative consequences of inaction, that's a problem to be solved once you have stopped actively digging a hole. That's a problem to be solved once you have stopped pushing people into that hole. That's a problem to be solved once you've stopped publicly congratulating people for not being pushed into the hole. By you. With an audience. There is nothing remotely subtle or nuanced about that distinction. There's an old adage that if you make laws people can't respect, then they will stop respecting the law. This is but one route to that problem.
- gerberthomas 4y agoReading the comments here, I see 2 things most of us seem to agree on: 1. team metrics are more useful than individual metrics; that makes sense because a team is an expression of a shared context while people come and go, so team metrics are inherently more valuable to the company in the long term 2. PR cycle time and some version of lead time (time to reach production) are often cited as 2 important metrics; also makes sense because those are the main critical, serial steps in software delivery Now, I would contend that the rest of the productivity metrics depend on what the team and its parent structures are trying to achieve, and should be DIFFERENT over teams and over time. Maybe a team with a spaghetti-like legacy project will want to track LOC or cyclomatic complexity for a while. Maybe some other team will want to track the amount of transitive dependencies. Maybe some will want to optimize for onboarding (time-to-10th-PR or something like this). 1 size will not fit all, and the team and its management must work to figure out what success looks like w.r.t productivity based on the context at hand.
- deleted 4y ago[deleted]
- roryokane 4y agoThis article written by Abi Noda in 2022 is suspiciously similar to the article https://www.usehaystack.io/blog/software-development-metrics-top-5-commonly-misused-metrics https://www.usehaystack.io/blog/software-development-metrics... written by Julian Colina in 2021. Not only do both articles talk about the same five metrics, listed in the same order, this newer article does nothing but paraphrase the original article in its description of each metric. For example, this newer article claims “a manager said … ‘Pull request count is the new vanity metric’”, while the original article stated that “Unfortunately it's a vanity metric”. And the last thing this newer article says about “impact” is that the name “sends a strong signal about how it should be used, particularly by managers”, matching the original article’s statement that “‘Impact’ suggests to executives and managers how this metrics should be used.” As the original article isn’t cited, this borders on plagiarism.
- geekjock 4y agoArticle author here — thanks for raising this. My article here was published in April 2021 and repurposed from my GitHub Universe talk given in 2019: https://www.youtube.com/watch?v=cRJZldsHS3c https://www.youtube.com/watch?v=cRJZldsHS3c Julian's article was published May 2021, one month after my Leaddev article, so hopefully it's pretty self-evident that Julian plagiarized my article pretty blatantly. I don't know Julian but I've heard from my former customers that his company has copied a lot of the work I did for my previous company, Pull Panda: https://github.blog/2019-06-17-github-acquires-pull-panda/ https://github.blog/2019-06-17-github-acquires-pull-panda/
- roryokane 4y agoOh, you’re right – your (Abi Noda’s) Leaddev article (https://leaddev.com/reporting-metrics/flawed-five-engineering-productivity-metrics https://leaddev.com/reporting-metrics/flawed-five-engineerin...) was published in 2021, not 2022. Sorry for getting that wrong. I think I confused the ad for “2022 Conferences & Events” at the start of the article with the publication date. It doesn’t seem true that Julian’s Haystack article (https://www.usehaystack.io/blog/software-development-metrics-top-5-commonly-misused-metrics https://www.usehaystack.io/blog/software-development-metrics...) was published in May 2021, even though that’s the date at the top of that article. It was apparently originally published in November 2020, as seen in a snapshot from December 2020: https://web.archive.org/web/20201204170826/https://www.usehaystack.io/blog/software-development-metrics-top-5-commonly-misused-metrics https://web.archive.org/web/20201204170826/https://www.useha.... Your article does seem to have been published in April 2021, as you say: https://web.archive.org/web/20210419071103/https://leaddev.com/reporting-metrics/flawed-five-engineering-productivity-metrics https://web.archive.org/web/20210419071103/https://leaddev.c.... Though your article was published later, the talk you gave on this topic (your YouTube link) was from December 2019, before either article. Julian Colina’s article quotes many phrases from that talk directly, so it does indeed seem that you, Abi Noda, are the original source.