3 ms·
I love that this post is quantifying results over actual data—a huge improvement over the speculative/anecdotal pattern that a lot of these analyses have. But
by blahedo 9y ago
I love that this post is quantifying results over actual data—a huge improvement over the speculative/anecdotal pattern that a lot of these analyses have.
But the conclusion that complexity and LOC both are higher in codebases with much unit testing... seems very weak. The complexity graph in particular looks very level, with the trendline driven almost entirely by the very low datum for 0-10%, and the LOC graph looks almost as suspicious. That low datum definitely demands further explanation before any conclusions are drawn on those two.
To me, it also inspires two further questions about the LOC methodology:
1) how sensitive is the measure to different coding styles? Does it include the method header? Does it include the closing bracket for the whole function? Does it include open bracket for the function when on a line by itself? Does it include any bracket on a line by itself? With the average method length ranging from 2 to 5, coding conventions on brackets could make a substantial difference, and if coding conventions correlate with testing philosophy at all (for whatever reason), that's a possible threat to validity.
2) Could these averages be dominated by a few larger projects? That is, is the average computed as
average (length of all methods, all projects)
or
average ( for each project, (average (length of methods)))
? If the former, larger projects would dominate the average. (True of all the other questions that involve counts and averages too, actually.)
- couchand 9y agoIn additions to the methodology concerns you raised, I have to question the use of a "linear" trend line, particularly when one axis is arbitrary buckets of unequal size.