4 ms·
I always say, somewhat some tongue-in-cheek, and somewhat intentionally provocatively, that if you can use stuff like TDD and pair programming, then you're prob
by kgo 16y ago
I always say, somewhat some tongue-in-cheek, and somewhat intentionally provocatively, that if you can use stuff like TDD and pair programming, then you're probably working on a boring problem.
And I think there's some truth to that. On a macro-level, how would you even begin to write tests for a search engine or some stock market bot or other notoriously hard problem?
search_on("avatar").should_return("http://www.imdb.com http://www.imdb.com)
best_stock_for(:percent_return, 200).should_return("cisco")
?
These problems an inherently non-deterministic. How do you even begin to write a test for that?
On a micro-level, sure, maybe you're working on a single component. And TDD would help you come up with the interface. But if you don't even know if the answer is a genetic algo, or simulated annealing, or using mechanical turk, or whatever, there's really no point in even trying to freeze the interface. Which is what TDD really does as much or even more than it verifies the resulting code. It defines the interface ahead of time. It's a way to trick developers into writing specifications without using that nasty imprecise context-sensitive language know as english.
But then again, right now we're rewriting a pretty critical piece of code. We've thought a lot about how it works. We had a few meetings about the new approach. Wrote up a quick email with a basic API. And doing pairing and TDD from there, well that's actually working out pretty well. And I'm confident we're getting better code quicker because of the approach.
Ultimately it gets back to the statement that real developers ship. In some cases, BDD and pairing will help you ship higher quality software quicker. In other cases it won't, and it'll end up wasting money and time. And real developers will then use their tools accordingly, and not dogmatically.
- nostrademons 16y agoGoogle has tests that read almost exactly like your examples. Of course, the whole search engine isn't specified by tests. However, if IMDB isn't in the top 10 results for [avatar] or barackobama.com isn't in the top 10 for [obama], something is seriously wrong and a human should look into it. The rest of your post is pretty good. Not sure why you've been downvoted.
- kgo 16y agoBut I imagine those are more like regression tests, right? Not write the test first BDD/TDD-style, then implement the whole search engine around it, tests. That was the implied context of that statement.
- bdclimber14 16y agoThese tests would send a red flag if they failed, but they aren't good as TDD tests. First, they are tied to a current time's context. Maybe next year, IMDB is no longer a good source for movie information, and it's so bad that it's on page 2. It could happen. In theory, your test cases should be consistent. Secondly, the TDD process is always to write code to pass tests. Well it's pretty easy to write code to return IMDB, but it's really hard to write test suite, that when coded against, would produce Google. That test would look like: search_on("avatar").should_return("www.imdb.com") if is_really_good_result(imdb) Once you are Google, then you should have these tests to make sure things keep working. However, I think it is really hard (in a bad way) to do TDD on hard problems.
- tooky 16y agoYou can't test for a particular outcome given an unknown set of input data. If you have a known snapshot of data, you can set expectations for how you want your system to behave under those circumstances. If you create a world where IMDB would rank highest for your indexing algorithm for the term "avatar", then you can expect that when you run a search it will be returned as the first result.
- bdclimber14 16y agoI once worked with some developers that literally did something close to what you described. In a nutshell, it was for an advanced job-employee recommendation algo. I knew how the algo should work, it was a form of collaborative filtering, but hadn't a clue how to write test cases. It was a really hard problem.
- ekidd 16y agoOn a macro-level, how would you even begin to write tests for a search engine or some stock market bot or other notoriously hard problem? These problems are hard because you evaluate your results with some kind of "quality score" instead of a simple "pass/fail" metric. You need to adapt your testing strategy accordingly. One useful strategy is to define a decent metric, and try to maximize it. Let's say that you have two sources of data: (1) A list of pages that should rank highly for specific queries, and (2) a list of links that users clicked on and stayed at without coming back to click on other links < 30 seconds later. Your goal is to write a search engine which ranks these links, on average, as highly as possible for the relevant queries. You do this with the usual techniques of experimental science: Split your data in half, develop against part of it, and hold a part in reserve for your final tests. If you make a clever change, and your average "good link" position drops from 1.7 to 3.9, you back your change out and try again. One of my clients actually did something similar with mapping software. Before deploying new code (or a new version of the map data), they ran extensive automated tests, and flagged anything weird for human review. I wrote them several test harnesses, one of which discovered a situation where the driving directions took a right hand turn off an overpass and tried to merge into traffic below... :-) Granted, these types of "tests" are essentially very high-level integration tests. But they serve the major purpose of tests: Automatically ferreting out disastrous mistakes before you ship to customers.
- kgo 16y agoI think this is another misunderstanding because I worded that question so poorly. Apologies. Yes splitting your data into a training set and a test set is a really good idea for problems like these. But I'd still argue it's near impossible to generate decent grading first, TDD-style, with only a superficial knowledge of the problem you're trying to solve in that case. Say you've got one obscure Star Wars character in your test set. Top results are probably imdb and wikipidia and that looks good. Then you've got a slightly less obscure character in the test set. And you stick with imdb and wikipedia as the top results for testing purposes. But he's got a huge page on wookiepedia. He's got a huge fansite. He's got a huge personal page. A twitter account more popular than Austin Kutcher. All popular enough to bump IMDB and Wikipedia out of the top five. Now the test you wrote is broken, and you're coding to broken results. In that case I don't think we're capable of generating good tests to define the result set FIRST, TDD-style. I don't think we're capable of specifying the application behavior FIRST, BDD-style. I think its better to play around with the data and algorithms first, until you start to get a gut feel for the data. And then you can write some decent tests for the test data. Same with mapping software. It's probably a good idea for the shipping product to have a test suite that does something like generate 100 random or not-so-random trips, then make sure that: you can get there from here, You can do it in +/- 5% of the time/miles that we've already esablished, etc. But I question how much value those tests have for day 1 or week one or even month one when you're writing the software. There's a lot of legwork before those will even come close to passing. And because you now don't have an un-broken build, people will potentially start ignoring problems with tests that should be passing week one. And I know you're not saying this, but every time people argue about TDD, the TDD proponents seem to think that no-one else tests. Which isn't true at all. You do need tests. The question is when and how. And the answer isn't always before anything else.
- felixge 16y ago> But if you don't even know if the answer is a genetic algo, or simulated annealing, or using mechanical turk, or whatever, there's really no point in even trying to freeze the interface. You are completely right. In this case you would start prototyping various solutions to the problem. Once you have explored the domain, you either go for implementing a certain approach using tdd, or you do more exploring. --fg