3 ms·
It's probably worth going a bit deeper into the paper before picking up conclusions. And I think the study could really do a bit of a better job of summarizing
by Delk 2y ago
It's probably worth going a bit deeper into the paper before picking up conclusions. And I think the study could really do a bit of a better job of summarizing its results.
The abstract and the conclusion only give a single percentage figure (26.08% increase in productivity, which probably has too many decimals) as the result. If you go a bit further, they give figures of 27 to 39 percent for juniors and 8 to 13 percent for seniors.
But if you go deeper, it looks like there's a lot of variation not only by seniority, but also by the company. Beside pull requests, results on the other outcome measures (commits, builds, build success rate) don't seem to be statistically significant at Microsoft, from what I can tell. And the PR increases only seem to be statistically significant for Microsoft, not for Accenture. And even then possibly only for juniors, but I'm not sure I can quite figure out if I've understood that correctly.
Of course the abstract and the conclusion have to summarize. But it really looks like the outcomes vary so much depending on the variables that I'm not sure it makes sense to give a single overall number even as a summary. Especially since statistical significance seems a bit hit-and-miss.
edit: better readability
- baxtr 2y agoRe 26.08%: I immediately question any study (outside of physics etc) that provides two decimal places.
- kkyr 2y agoWhy?
- jayd16 2y agoThe assumption is that they almost certainly did not have the sample size to justify two decimal places. If they want two decimals for aesthetics over properly honoring significant figures then it calls the scientific rigor into question.
- energy123 2y agoBut they say "26.08% increase (SE: 10.3%)", so they make it clear that there's a lot of uncertainty around that number. They could have said "26% (rounded to 0 dp)" or something, but that conveys even less information about the amount of uncertainty than just saying what the standard error is.
- Delk 2y agoThey could still have gone for one decimal. Or possibly even none, considering the magnitude of the SE, but I get that they might not want to say "SE: 10%". The second decimal point doesn't essentially add any information because the data can't provide information at that precision. It's noise but being included in the summary result makes it implicitly look like information. Which is exactly why including it seems a bit questionable. That's not the major issue with the study, though, it's just one of the things that caught my eye originally.
- svnt 2y agoTo get a better picture of how this comes about: Microsoft has a study for their own internal product use, and wants to show its efficacy. The results aren’t as broadly successful as one would hope. Accenture is the kind of company that cooperates and co-markets with large orgs like Microsoft. With ~300 devs in the pool they hardly move the population at all, and they cannot be assumed to be objective since they are building a marketing/consulting division around AI workflows. The third anonymous company didn’t actually have a randomized controlled trial, so it is difficult to say how one should combine their results with the RCTs. Additionally, I am sure that more than one large technology company went through similar trials and were interested in knowing the efficacy of them. That is to say, we can assume other data exist than just those included in the results. Why did they select these companies, from a larger sample set? Probably because Microsoft and Accenture are incentivized by adoption, and this third company was picked through p-hacking. In particular, this statement in the abstract is a very bad sign: > Though each separate experiment is noisy, combined across all three experiments It is essentially an admission that individually, the companies don’t have statistically significant results, but when we combine these three (and probably only these three) populations we get significant results. This is not science.
- Delk 2y agoYeah, it does seem fishy. The third company seems a bit weird to include in other ways as well. In raw numbers in table 1, there seem to be exactly zero effects from the use of CoPilot. Through the use of their regression model -- which introduces other predictors such as developer-fixed and week-fixed effects -- they somehow get an estimated effect of +54%(!) from CoPilot in the number of PRs. But the standard deviations are so far through the roof that the +54% is statistically insignificant within the population of 3000 devs. Also, they explain the introduction of the week fixed effect as a means of controlling for holidays etc., but to me it sounds like it could also introduce a lot of unwarranted flexibility into the model. But this is a part where I don't understand their methodology well enough to tell whether that's a problem or not. I generally err towards the benefit of the doubt when I don't fully understand or know something, which is why I focused more on the presentation of the results than on criticizing the study and its methodology in general. I'd have been okay with the summary saying "we got an increase of 27.3% for Microsoft and no statistically significant results for other participants". But perhaps I should have been more critical.
- throwthrowuknow 2y agoPersonally, my feeling is that some of the difference is accounted for by senior devs who are applying their experience in code review and testing to the generated code and are therefore spending more time pushing back asking for changes, rejecting bad generations and taking time to implement tests to ensure the new code or refactor works as expected. The junior devs are seeing more throughput because they are working on tasks that are easier for the LLM to do right or making the mistake of accepting the first draft because it LGTM. There is skill involved in using generative code models and it’s the same skill you need for delegating work to others and integrating solutions from multiple authors into a cohesive system.