7 ms·
How do you approach this issue? http://www.distilled.net/blog/conversion-rate-optimization/why-your-cro-tests-fail/ http://www.distilled.net/blog/conversion-ra
by will_critchlow 14y ago
How do you approach this issue?
http://www.distilled.net/blog/conversion-rate-optimization/why-your-cro-tests-fail/ http://www.distilled.net/blog/conversion-rate-optimization/w...
I haven't come up with a decent answer yet especially for smaller sites that can't run tests on very specific segments...
- btilly 14y agoBased on your article and code, I believe that you have misunderstood what the statistical test is supposed to be telling you. When you make decisions at 95% confidence, your guarantee is that at most 5% of the time are you going to wrongly conclude that one is better when it isn't. However you have absolutely no guarantees about having correctly called the direction of the test if you do hit that significance level. (Indeed if the null hypothesis is correct, every time you call the test, you're wrong!) What you did is simulated the test many times, ignored the cases where you were told there was no answer (thereby throwing away a large part of your guarantee), and found that you could be wrong a large portion of the time. Furthermore if there was a large, discoverable, random factor, you found that could be correlated with a lot of the mistakes. Unfortunately, in addition to the discoverable factor, there are lots of unknown random factors that also get randomly correlated. And even if there aren't, there is always the possibility of experiencing bad luck. The only way to solve this is to throw enough traffic at the problem that the underlying bias can be reliably detected statistically. There is a complicated relationship between the size of bias you're willing to get wrong, and the amount of data that you need to collect to reliably detect it. I hope to explore that relationship in the next two articles. For larger sites, that solution is perfectly workable. For smaller ones it is not, and the best that I can suggest is that they rely heavily on design principles that have been validated through A/B testing on larger sites, and hope they are not going too far wrong.
- will_critchlow 14y agoI'm pretty sure we should just ignore my code for the purposes of this discussion - my math(s) may be fuzzy, but my code has never been a strong point. I was mobile when I wrote the original question - perhaps a better way of phrasing it would have been something like: Don't many (most? all?) of these theoretical approaches assume that the sequence of results for each page (call them a_i and b_i for i=1,2,3... where each x_n is 0 [no conversion] or 1 [conversion]) are sequences of iid random variables with underlying conversion probabilities p_a and p_b? In reality, these sequences are much more complex and, if the scale of the variation in conversion probability within the sequence is greater than the difference between p_a and p_b won't the test be much weaker than we originally thought? To use an example that is simplified vs. reality but hopefully indicates what I mean, imagine that we have two traffic sources - one with a conversion probability twice that of the other (on each page variant - so p_1_a = 2p_2_a and p_1_b = 2p_2_b). We randomly send traffic from both sources (1 and 2) to each variant (a and b). Do the standard tests work even though our sequence of conversions are not iid?
- btilly 14y agoNo. One good way to see why is this. The theoretical approaches that you're talking about do not care about the internal details of your random number generator. You could have p_a be the result of a single random decision (convert or not), or be the result of first finding out that we had a random source, then having conversion probabilities that depend on the source. Either way, as long as in the end you get a stable p_a of converting from noticing you have a new visitor to an actual conversion, probability theory says that the exact same statistical statements will be true, and you'll have entirely equivalent results. (Note, the way I set up my particular approach means that actual conversion rates can shift over the test without invalidating the results. I didn't call that out too strongly, but I value that detail.) However there is a gotcha lurking. The gotcha is that if conversion depends on your source, then A/B testing will optimize for your current traffic mix, but its result may become wrong if that mix shifts. My usual approach is to simply assume that the world is not going to be malicious in this way, unless I have specific reason to suspect it is. Thus, for instance, I would not think twice about source vs button color. But if I had a landing page with random testimonials, I'd expect that a testimonial from pg would convert better than a testimonial from Phil Ivey. And conversely for traffic from a gambling site.