8 ms·
TDD, BDD, and the tea tasting lady
- Sandman 14y agoI'd be genuinely interested in reading a paper on what impact these methodologies have on software development. Does anybody know of any research on this topic?
- rdfi 14y agoSince I wrote that article I've found one or two papers about the subject, one of them was one that I heard about in .Net Rocks, it compared code reviews with unit testing, but I can't find it :S While I was searching I found this one from NCSU and Microsoft: "On the Effectiveness of Unit Test Automation at Microsoft": http://collaboration.csc.ncsu.edu/laurie/Papers/Unit_testing_cameraReady.pdf http://collaboration.csc.ncsu.edu/laurie/Papers/Unit_testing... It's not about TDD although they mention it in the paper. I'll try to find that one about code reviews and post it here as well
- rdfi 14y agoFound that paper about code reviews, it is mentioned in this blog post: http://kev.inburke.com/kevin/the-best-ways-to-find-bugs-in-your-code/ http://kev.inburke.com/kevin/the-best-ways-to-find-bugs-in-y... And here's the link to the paper itself: http://kev.inburke.com/docs/basili_testing.pdf http://kev.inburke.com/docs/basili_testing.pdf
- raverbashing 14y agoAll those practices are good (tdd, code review, etc) if people know what they're doing Otherwise: Sometimes code review will turn into a pissing match of who can find the most errors in other people's code TDD will turn into a quest for a higher "coverage number" regardless of the actual testing quality
- rdfi 14y agoA pissing match of who can find the most errors in other people's code is probably a good thing :P It's when it starts being about personal aesthetic preferences without any consequences in terms of code readability/quality that I think it starts to be annoying.
- raverbashing 14y agoYes, there's usually a finite amount of errors in a given code. But it's mostly about, as you said, personal preferences or 'philosophies' regardless of the actual results of the code. So you waste a lot of time because someone thinks the way you did it is not 'OO enough' or 'should be organized better' even if it is working (I'm not talking about code that's messy)
- e12e 14y agoI found this: http://proceedings.informingscience.org/InSITE2012/InSITE12p165-187Bulajic0052.pdf http://proceedings.informingscience.org/InSITE2012/InSITE12p... Among other things, it includes a long list of references -- taken after searching based on: http://mango2.vtt.fi/virtual/agile/publications.html http://mango2.vtt.fi/virtual/agile/publications.html That again was referenced in the following blog post from 2007, that I found more interesting (if rather similar) to op: http://www.artima.com/weblogs/viewpost.jsp?thread=216434 http://www.artima.com/weblogs/viewpost.jsp?thread=216434
- e12e 14y agoI also came across this interview/discussion: "Jim Coplien and Bob Martin Debate TDD" http://www.youtube.com/watch?v=KtHQGs3zFAM http://www.youtube.com/watch?v=KtHQGs3zFAM
- stonemetal 14y agohttp://evidencebasedse.com/ http://evidencebasedse.com/ Collects papers on SE practices. They claim to have 33 papers on TDD and 40 on testing in general.
- rdfi 14y agoThanks for this
- lgunsch 14y agoI found this paper NCSU a while back: http://staff.unak.is/andy/MScTestingMaintenance/Homeworks/STMHeima7TestDrivenDevelopment.pdf http://staff.unak.is/andy/MScTestingMaintenance/Homeworks/ST...
- mistercow 14y ago>Your reaction (as mine was) is that it is impossible to tell the difference. If you pour the milk into the tea, the first bit of milk will heat to nearly boiling, scalding it and changing its flavor. If you pour the tea into the milk, you don't have that problem.
- skore 14y agoAnd, according to some research[0], no matter which way you pour it, milk will make the tea worse. (And as a personal note - if you need milk to make your tea taste good, maybe you just don't have good tea?) [0] http://news.bbc.co.uk/2/hi/6241139.stm http://news.bbc.co.uk/2/hi/6241139.stm
- mistercow 14y agoThe health problem can be solved by using a non-dairy milk like almond milk (which is also objectively tastier). As far as snootiness, I really think it's just a matter of personal preference. I enjoy it either way, although it can certainly mask the flavor of bad tea.
- waivej 14y agoI drink watered down orange juice. It mixes evenly if you pour the water first. Otherwise the concentration is higher and the temperature is lower near the bottom of the glass.
- mistercow 14y agoI'd hypothesize that that's a different effect. Orange juice is denser than water, so if you pour water into orange juice, a lot of it will float on top. If you pour orange juice into water, on the other hand, the orange juice will sink through the water, mixing relatively evenly. By the same reasoning, tea should actually mix better if you add the tea first, since milk has about the same density as orange juice and tea has about the same density as water. But in the case of tea and milk, you just have to bite the bullet and stir it manually if you want to avoid scalding.
- 14y ago
- ajanuary 14y agoThe article talks about how to blind the tea tasting lady to make the results more accurate. It's pretty hard to apply that same blinding to reading code with large or small methods. There are definitely ways to test it, but it's not quite as simple.
- evolve2k 14y agoThere is a great book which teaches you statistics looking through its fascinating history. The Lady Tasting Tea: How Statistics Revolutionized Science in the Twentieth Century http://www.goodreads.com/book/show/106350.The_Lady_Tasting_Tea http://www.goodreads.com/book/show/106350.The_Lady_Tasting_T...
- swalsh 14y agoIt's perhaps not rigorously scientific, but our red mine server is tied into our Jenkins continuous integration server, which also has a code coverage tool tied in. We've been doing this for a little while, and have the ability to visualize bug counts divided by severity along side test count, and code coverage percentage. There is an order of magnitude difference from the projects that did not use TDD, and the projects that did. Another nice part of red mine is that I can also visualize how close we are at meeting our expectations for time estimates (they have this task tracking feature). Overall adding TDD has made development a bit slower in the front end, but a lot shorter in the back end (during QA). However when you look at the past, it would seem like something like 60% to 70% of our time was actually spent fixing issues found in QA, which was usually under estimated. So overall we're definitely seeing quantitative proof that TDD is a better methodology in comparison to our past process. Anther cool tool for project management is we've built a set of tools that allow us to designate a requirement from our functional spec with a unit test, which runs every check in. So our project manager can get a near real time assessment of our progress. I also use that report and associate it with the code coverage report. Though this part of the system is newer, so I have no data on how effective it is.
- jbrechtel 14y agoIt'd be great if you would write a blog post with more details about the TDD vs non-TDD projects. Things like: Duration of observation team size (over time) survey of team members anyone go from one team to another? Someone with perspective across both teams would be good. Domain complexity of each project Other code metrics (cyclomatic complexity, coupling, etc) A post would be nice as a reference (as opposed to a comment here). Not a small request, I understand, but I'm sure others would find it very valuable too.
- coopdog 14y agoDo you find linking unit tests to functional requirements useful? I've also been toying with visual traceability (requirementweaver.com) and it seems to have a lot of potential to speed up the traditionally slow 'quality process'
- 14y ago
- abraininavat 14y agoI think that Uncle Bob and his ilk would dispute the idea that breaking a large function into smaller functions ever makes the original function less readable. They'd argue that each of the new, small functions is proven to be correct (since you'd fully unit-tested it), so its contents are of no consequence when reading the larger function. How true that is I'm not really sure.
- jbrechtel 14y agoIn an OO-world function extraction usually results in private methods to a class. You (generally) don't test those directly so it doesn't necessarily result in more specific unit tests.
- maio 14y agoWell you can extract methods into new object where they will be public.
- chadcf 14y agoI think you have to have a good balance here. I've found following a huge method or function rather complicated, but at the same time I think I find trying to track down complicated abstractions even more complicated. If you've ever stared at a new codebase trying to track down a bug, you know the joy of grepping and following chains of methods trying to get to the part that actually matters. This may not make it technically less readable but I'd say it makes it significantly less, uh, understandable? Really, it boils down to how easy the code is to understand and maintain. Group code into logical groups. Don't break it up if the only benefit is following some arbitrary rule stating your functions should be no more than x lines long.
- rdfi 14y agoI believe the rule is not necessarily about having a small number of lines (that is just a consequence), it's about the method having one responsibility, only do one thing, and that thing should be understandable just by reading the method's name. The ultimate goal is that the code can almost be read as plain English. This has the consequence of making the methods smaller since they are narrow in scope. It's easier to understand a well named method with a few lines of code (and believe me, it's easier to properly name it than a method with a lot of code, since the former is focused in one task and it is easy to come up with a name that describes that task; and the latter, where the method does so many things that there's no way you can name it properly [have you ever found methods with names like DoWork and then 100 lines, I'd say that is a code-smell].
- noelwelsh 14y agoThe problem is one of cost. The tea tasting experiment is very easy to reproduce and very cheap to run. Reproducing, say, the development process of Myna is practically impossible not to mention the prohibitive cost of finding devs as awesome as Dave and I ;-) To apply the scientific method to software development you need to apply methods from the social sciences. Now I don't know a great deal about these, but I do know you need lots of data which is very hard to find. The typical solution is test hypotheses on undergrad students, because that is what the experimenters have plentiful access to. The problem, which is also apparent in psychology, is generalising these results beyond this group. Are the experiences of 2nd year undergrads using Java for a 2 week project predictive of developers with 10+ years experience working on a year long project? One can reasonably argue they are not.
- rlpb 14y agoIt's very difficult to perform a fair test. If I test random programmers by getting them to do the same project either with TDD or without, TDD proponents might (justifiably) complain that the programmers didn't do TDD properly, since they didn't know how. If I use TDD proponents to do the TDD side of the test, and TDD skeptics the non-TDD side, then one might claim that there is a bias in the quality of programmer (either people who like TDD make better programmers, or TDD skeptics make better programmers, depending on the results). For every test you come up with, there's a potential bias that you cannot eliminate, since there will be some hypothetical correlation somewhere that doesn't correspond to TDD itself.
- JohnLBevan 14y agoDo four tests - TDDers doing TDD, TDDers doing non-TDD, non-TDDers doing TDD and non-TDDers doing non-TDD. From this you can get a better idea of whether it's the method or the coder's skill affecting the results. Clearly if TDD is better the TDDers will do better than the non-TDDers when both doing TDD because they've more experience. Comparing the non-TDDers in TDD vs non-TDD will show you if the method alone is enough, or if other factors have affected it (e.g. learning curve of the new method, ability of programmers who chose that method). The TDDers doing non-TDD helps contribute to the question of are they just better programmers (i.e. do they do comparatively well in non-TDD / are they worse that those TDDers doing TDD). Doing other tests with entirely randomly chosen programmers ensuring they're trained up in these techniques would be another way - though as you point out the training would have to ensure they got the techniques they were being tested against / you'd need to avoid bias of people having spent the last few weeks exclusively in an exclusively TDD environment due to the training.
- JohnLBevan 14y agoAnother thing to take into account is who's coding, what they're coding in, what they're building, how big their team is, how mature the team is (i.e. how long have they been working together) and how their environment's set up. As with most questions, context is everything, so should anyone experiment to find the best solution you'd want them to include those details in their conclusions, and ideally to alter and test against each of those variables and show the effects of BDD/TDD under those conditions.
- waivej 14y agoI've been thinking about using a random process to guide bug seeding. In other words, putting in bugs in random locations and random "type" and then see how many are caught by the test suite.
- colonelxc 14y agoThis is called mutation testing. Here is an article I read recently that I thought was a good introduction to the idea: http://dev.theladders.com/2013/02/mutation-testing-with-pit-a-step-beyond-normal-code-coverage/ http://dev.theladders.com/2013/02/mutation-testing-with-pit-...