4 ms·
Regarding coverage comment, your right in that code can contain capabilities (or bugs) beyond the TDD coded conceived. (If they weren't conceived then it could
by lancerkind 9y ago
Regarding coverage comment, your right in that code can contain capabilities (or bugs) beyond the TDD coded conceived. (If they weren't conceived then it could be a bug or things that don't happen in practical operations.) The TDD people are right in that if you follow the TDD process, all your code comes up as covered by a code coverage tool. But this is only coverage. Not all possible functionality coverage. I imagine "all possible functionality coverage" is an intractable problem in the same class as the Halting Problem. Branch/line coverage should be near 100% for TDD code. (Some I/O coupling areas are intentionally saved for other types testing as they aren't unit testable.) It's an easy thing to measure. It's a common industry metric. It's not the whole picture but it is something.
Measuring code coverage on TDD created software that was built without monitoring the code coverage, then later use code coverage at the end as an evaluation: I've only a few anecdotal (two) instances where I did this and I found the code coverage to be well above 95% for the code base and at 100% for modules that were pure logic. In one of those instances, after many rounds of defect injection, I discovered one or two locations where another unit test could be added.
All in all, I've been pretty happy with TDD. But who cares? I want to hear from the industry why others aren't doing it. :-)
Thank you eesmith, I enjoyed the discussion!
How about hearing some more from others? What's stopping you from doing TDD? Is it because you don't believe it? If you don't believe it, how was it presented to you? Did you try it?
- eesmith 9y ago"your right in that code can contain capabilities (or bugs) beyond the TDD coded conceived" I do not understand what you mean by my being right. I did not say that. As you wrote it, it's a trivially true statement, and I didn't think what I wrote was so trite. I will try to reword it. In the "Red - Green - Refactor" style of TDD you are allowed to refactor the code so long as the tests remain green. This assumes there are sufficient tests to catch errors in the refactoring. However, one of the allowed refactorings is "Substitute Algorithm"[0]. The new algorithm may have have corner cases which the old algorithm did not have, and so wasn't originally tested. R-G-R omits the part where you add new unit tests which are expected to be green, but added in order to verify that the new algorithm really does work correctly for new corner cases. It should be "Red-Green-Refactor-Yellow-Green", where "Yellow" means "do the failure analysis to ensure that the refactoring didn't add new failure cases." Yet TDD people don't teach that. And I think the reason why is that this "Yellow" stage is the normal part of non-TDD programming. If you drop "Red-Green" and just iterate over "Code-Green-Yellow-Green" (write the code, make sure the old tests work, do the failure analysis, and make the new tests work), you'll still end up with good quality code. This comes back to another aspect of TDD. Suppose that the programmers in a company are in their 20s and 30s while middle management is in their 50s. There is an asymmetry as middle management has much more experience in applying pressure than the younger programmers. Managers, after all, have more experience trying to manipulate people. Some programmers in this situation use TDD as a tactic to resist external pressure to release code with insufficient testing. "The code's done, but it isn't fully tested." "It's done? Ship it, and we'll deal with the problems later." With TDD this becomes "we're using TDD because <famous people> say it's a best industry practice". This gives the programmers a way to externalize the resistance to releasing poorly tested software. However, it's also unprofessional. Some code can be lightly tested, perhaps manually tested instead of a unit test suite. TDD shuts off the dialog about how to make those trade-offs at the business level, in the name of either technical purity or an unwillingness or inability to really negotiate. "if you follow the TDD process, all your code comes up as covered by a code coverage tool" No, that is incorrect. I have given two examples of how TDD code can end up with less than 100% statement coverage, much less 100% branch coverage. Ken Beck wrote "TDD followed religiously should result in 100% statement coverage." You backed that off a little with "should be near 100%". But my point is that these are assertions of faith, which TDD proponents repeat without evidence. If they repeat this without evidence, what else are they repeating without evidence? I myself have written non-TDD code, then tests, then verified code coverage, and ended with 98% coverage. (I wasn't able to test malloc() failures, though I coded in handlers for them and manually exercised those code paths to make sure they worked.) [0] As Fowler originally described it, "Substitute Algorithm" is used "to replace an algorithm with one that is clearer." However, if a new test appears to hang, and analysis shows it will take several days to finish, then you may need to use a more complex algorithm with better scaling performance. If this isn't a refactoring, then what is it? Others, like https://refactoring.guru/substitute-algorithm https://refactoring.guru/substitute-algorithm , don't require that the new algorithm be simpler.
- lancerkind 9y agoI'm interested in your test later approach. When you're attempting 100% branch coverage, do you sometimes find writing the test later causes redesign of the product code? On average, how many days does it take to complete a cycle of: write the code and then write automated unit test code? What ratio of that time is spent doing test automation versus writing production code?
- eesmith 9y agoI rarely try for 100% branch coverage. I find that the extra tests aren't worthwhile. I don't even try for 100% statement coverage. I use coverage to help identify important missing test cases. For example, some of my code will read/write a binary format of my own devising. The format is used in a "friendly" environment, so I don't have to worry about malicious users trying to cause failures. (The code is in Python, which eliminates a lot of the memory error that a malicious user could use.) I could get 100% coverage by creating corrupted binary files in all of the different ways to fail. But why go through the work of testing unrealistic cases? Instead, I use my experience and intuition to only test the realistic cases. I also know that my code is not mission critical, so if one of my users reports a failure, I'll send them a patch or point release. I can't give you a breakdown of time because I don't track it. My usual estimate is 1/3 programming, 1/3 testing, and 1/3 documentation, for commercial-quality packages. I find that writing good documentation helps find bugs that users are likely to experience. (Now that I think of it, that binary format took several months to write. The first two attempts didn't scale and I had to toss them. They had almost no tests, so think of them as spikes.) For consulting, it depends on the project. In one two-week project I had no tests. The control flow was simple and most of the code was executed during normal use. Plus, the support staff was able to maintain the software in case of errors, and we did a thorough code review for knowledge transfer. In another project, I was brought in to optimize some existing code, which didn't have tests but which had been extensively tested manually. It was presumed to be correct. The new code I wrote also didn't have unit tests. Instead, I cross-compared a few million inputs to check that they gave the same results, at about 10x performance. In yet another project, I had unit tests for the user-facing parts, but not for the main algorithm. Rather than explain the program, I'll use an analogy. Consider a program to factor a number into its primes. This is difficult. On the other hand, it's easy to test that the result is correct. Multiply the numbers to see if they give the original number, and use a primality test to ensure all of the factors are prime. In essence, I have a #define which enables verification. This also slows things down, so it's not enabled by default. I then (once again) processed a few million inputs. This sort of integration testing is much more complete than manual unit tests. (I do have a simple unit test used as a burn test, so it's not completely untested.) For an extreme case, see https://randomascii.wordpress.com/2014/01/27/theres-only-four-billion-floatsso-test-them-all/ https://randomascii.wordpress.com/2014/01/27/theres-only-fou... . Why write unit tests when it's possible to examine all 2^32 possibilities in 90 seconds?