5 ms·
I'm pretty sure this formula is correct, but I haven't seen it published anywhere. John Cook has some veiled references to a closed-form solution when one of th
by EvanMiller 12y ago
I'm pretty sure this formula is correct, but I haven't seen it published anywhere. John Cook has some veiled references to a closed-form solution when one of the parameters is an integer:
http://www.mdanderson.org/education-and-research/departments-programs-and-labs/departments-and-divisions/division-of-quantitative-sciences/research/biostats-utmdabtr-005-05.pdf http://www.mdanderson.org/education-and-research/departments... [pdf]
But he doesn't really say what that closed form is, so I think his version must have been pretty hairy. (My version requires all four parameters to be integers, so I doubt we were talking about the same thing.)
Sadly I couldn't get the math to work out for producing a confidence interval on |p_B - p_A| so for now you're stuck with Monte Carlo for confidence bounds.
Thanks to prodding from Steven Noble over at Stripe, I'll have another formula up soon for asking the same question using count data instead of binary data. Stay tuned!
- octernion 12y agoThanks for doing this, Evan! Much appreciated - I've already implemented it to compare with our traditional confidence score based A/B metrics. I'm not sure what you mean by Monte Carlo for confidence bounds - is there some reading I can do to brush up on that? edit: here is your solution implemented in Ruby as I implemented it https://gist.github.com/nickelser/6dfa0f2737c502dae267 https://gist.github.com/nickelser/6dfa0f2737c502dae267
- paraschopra 12y agoConfidence interval means that is expected conversion rate of A is 0.5 and for B is 0.4, what is the actual spread. In other words, say, for 95% of the time does the conversion rate fall in the 0.3-0.5 range or 0.35-0.45 range. More data you collect , narrower this range will be. Confidence bounds for individual conversion rates can be calculated from the formula since they follow beta distribution. What Evan is saying that the confidence bound for the difference in conversion rates would need to be calculated through repeated simulations and taking (say) 5 percentile and 95 percentile of the values that come. For example, given you know probabilities for A and B (say 0.5 and 0.4) and that number of trials has been 100, you can simply run thousands of simulations, know what the conversion rates come out to be in respective simulations and aggregate the difference in conversion rates at the end of each simulation. That would be a confidence bound.
- octernion 12y agoGotcha - thanks so much for writing this up. I'll see about actually implementing this :)
- grayclhn 12y agoI don't do this stuff, so this may be a stupid question. But it looks like you have p_A and p_B being independent in the posterior distribution. How does that work? I would assume that they're pretty correlated IRL --- we're not changing that much in A/B testing. This must be coming from the likelihood function, so you may want to be more specific about the exact experiment you have in mind. I'm not sure off the top of my head whether imposing independence is going to overstate or understate the precision of the estimate...
- noelwelsh 12y agoThis paper by John Cook has the inequalities: http://biostats.bepress.com/mdandersonbiostat/paper54/ http://biostats.bepress.com/mdandersonbiostat/paper54/
- ajtulloch 12y agoThe key background here that is somewhat brushed over in the post is that the Beta distribution is conjugate to the Binomial - in the sense that if your parameter of interest theta has a Beta(a, b) prior, and you observe n successes and r failures from a Binomial trial with parameter(n, theta), your posterior is now Beta(a + n, b + r). The 'non-informative' distribution can be interpreted as having no special information about the prior information of theta, so you just take theta uniform over [0, 1]. This is identical to the Beta(1, 1) distribution. Thus, you obtain the update equations @EvanMiller derived, while making it more clear how to generalize to the informative prior case. In general though, I would have thought one would care more about distributional properties of the posterior p_B - p_A, which are more easily computed by MC methods - the closed form solution for P(p_B > p_A) is nice and fast, but it's not like hypothesis testing is in the critical path of a program, so I wouldn't shy away from MC methods here for a more nuanced view of your posteriors.
- sesqu 12y agoWhere do you use the fact that all the parameters are integers? I can only see the requirement for the one parameter you sum over.