4 ms·
A/B Testing is not Snake Oil.
- chrisaycock 16y agoIn the spirit of A/B testing, perhaps you could point to some reasonable and well-planned tests that produced no actionable results. Contrast that against a test that did produce actionable results. Then compare what went wrong and what, if any, lessons can be extrapolated. That would make for an interesting---and possibly meta---blog post.
- paraschopra 16y agoYep, that's a great idea. In fact, I had blogged about one in the past. Here it is: http://visualwebsiteoptimizer.com/split-testing-blog/left-vs-right-sidebar-which-layout-works-best/ http://visualwebsiteoptimizer.com/split-testing-blog/left-vs...
- elbrodeur 16y agoA/B testing is just one component of a good product development cycle. The easiest pitfall, though, is that the data only tells you what -- not why. This can lead you to make data-based design decisions blindly. I think the ideal design and development cycle incorporates A/B testing at the end to spot any outliers: You should already have done in-person usability and user testing. It's extremely cheap and, in some cases, almost magical. Seeing a real person use your product will guarantee you plenty of "a-ha!" moments. After testing with real users, I think it's appropriate to test with statistical users: You've already tackled the most glaring issues with what you're building, and the knowledge of how real people use your website can often help show you the "why" of the data, rather than just "what". Steve Krug's "Don't Make Me Think" has a chapter near the end on user testing on a shoestring budget. I'd highly recommend it.
- Jabbles 16y agoIf you run 100 different A/B tests and only 1 of them produces good results, you are only going to publish about that one good result and not the 99 other unsuccessful results. You'd better be sure those "1 in a 100" results had a confidence level well above 99%. I think misunderstandings of statistical analysis play a large part in the mistrust some people (mistakenly) have in A/B testing. If you get your analysis wrong, you'll slowly realise that the promised gains of the test turn out to be lies, hence the comparison to snake oil.
- paraschopra 16y agoWe only publish results that are statistically significant (albeit with 95%+ confidence)
- Jabbles 16y agoBut... that's my point. If you run 20 completely random tests, the probability that at least one of them gives you a "95% confidence limit" is ~63%. Statistical significance has to be adjusted depending on the number of times you run your test, else it loses all its value. I've already posted this link today, but I really recommend this guide: http://www.evanmiller.org/how-not-to-run-an-ab-test.html http://www.evanmiller.org/how-not-to-run-an-ab-test.html
- paraschopra 16y agoIt is 100 different tests, not same test run 100 times.
- neild 16y agoThat doesn't affect Jabbles's point in any way. Run enough tests, and you will get statistically significant but bogus results.
- paraschopra 16y agoOf course, that's the whole point of declaring statistical significance. With 95%, you will have false positives. With 99%, you will have false positives. There are no guarantees. Many times, when we declare a winner in test and if you keep running it, you may see that eventually it does not perform as good. But I agree my comment doesn't negate Jabbles' point.
- Jabbles 16y agoAh, but do you understand that the chance of a false positive can be much higher than 5% even with your "95% confidence limit"?