6 ms·
I'm pretty sure we should just ignore my code for the purposes of this discussion - my math(s) may be fuzzy, but my code has never been a strong point. I was m
by will_critchlow 14y ago
I'm pretty sure we should just ignore my code for the purposes of this discussion - my math(s) may be fuzzy, but my code has never been a strong point.
I was mobile when I wrote the original question - perhaps a better way of phrasing it would have been something like:
Don't many (most? all?) of these theoretical approaches assume that the sequence of results for each page (call them a_i and b_i for i=1,2,3... where each x_n is 0 [no conversion] or 1 [conversion]) are sequences of iid random variables with underlying conversion probabilities p_a and p_b? In reality, these sequences are much more complex and, if the scale of the variation in conversion probability within the sequence is greater than the difference between p_a and p_b won't the test be much weaker than we originally thought?
To use an example that is simplified vs. reality but hopefully indicates what I mean, imagine that we have two traffic sources - one with a conversion probability twice that of the other (on each page variant - so p_1_a = 2p_2_a and p_1_b = 2p_2_b). We randomly send traffic from both sources (1 and 2) to each variant (a and b). Do the standard tests work even though our sequence of conversions are not iid?
- btilly 14y agoNo. One good way to see why is this. The theoretical approaches that you're talking about do not care about the internal details of your random number generator. You could have p_a be the result of a single random decision (convert or not), or be the result of first finding out that we had a random source, then having conversion probabilities that depend on the source. Either way, as long as in the end you get a stable p_a of converting from noticing you have a new visitor to an actual conversion, probability theory says that the exact same statistical statements will be true, and you'll have entirely equivalent results. (Note, the way I set up my particular approach means that actual conversion rates can shift over the test without invalidating the results. I didn't call that out too strongly, but I value that detail.) However there is a gotcha lurking. The gotcha is that if conversion depends on your source, then A/B testing will optimize for your current traffic mix, but its result may become wrong if that mix shifts. My usual approach is to simply assume that the world is not going to be malicious in this way, unless I have specific reason to suspect it is. Thus, for instance, I would not think twice about source vs button color. But if I had a landing page with random testimonials, I'd expect that a testimonial from pg would convert better than a testimonial from Phil Ivey. And conversely for traffic from a gambling site.