4 ms·
There are two parts: 1) what is an appropriate statistical model for A/B testing and 2) how should we make decisions based on our current beliefs (the Bayesian
by equark 15y ago
There are two parts: 1) what is an appropriate statistical model for A/B testing and 2) how should we make decisions based on our current beliefs (the Bayesian posterior).
A sensible starting point for the first is a hierarchical beta-binomial model. For instance:
http://www.stat.cmu.edu/~brian/724/week06/lec15-mcmc2.pdf http://www.stat.cmu.edu/~brian/724/week06/lec15-mcmc2.pdf
Translating that example, the binomial variable represents the number of conversions given the total number of exposures. So if you show a red button 100 times and 10 people convert, then, using the notation in that PDF, n_i = 100 and y_i=10. We are interested in p(\theta_i|y_i, n_i), the posterior distribution of the conversion rate for experiment variation i (red, blue, green) given our data.
The hierarchical part of the model is what's Bayesian. Here we use a Beta prior, since \theta_i is between 0 and 1. This prior shrinks each estimate towards to overall conversion rate based on how much variation there is between experiments -- the \alpha and \beta parameters. You can think of \alpha and \beta as pseudo-observations -- the number of conversions and failures you've "seen" apriori. Given that we have multiple experiments, you actually have a sense for the distribution of \theta_i, and we can therefore estimate \alpha and \beta by adding a third layer p(\alpha, \beta).
There are many ways to make a richer model, but if you haven't seen Bayesian modeling before that's probably enough.
The beauty of the Bayesian approach is the posterior is what you want -- your belief about conversion given the data you observe, the model you assume, and your prior beliefs. As you add data, your posterior beliefs update, but at every point in time it always represents your best guess.
It solves the multiple comparison problem via shrinkage rather than by adjusting p-values. This is intuitive. If you see an outlier and you don't have much data yet, then it's probably just a random fluctuation and your prior shrinks your best guess towards what you think conversion rates should be overall. For instance, if you believe conversion rates are typically .05 and never .2, then if you see something like .2 after just a few observations, you'll probably guess the true \theta_i is more like .08.
The second part of the problem, optimal sequential decision-making is more tricky. It's a bandit problem, where there's a tradeoff between exploration and exploitation. As far as I'm aware, this is still considered a very tricky problem to solve optimally in all put the most simple cases. Practically you could probably get close to the optimal answer via forward simulation. There's a lot written on Bayesian bandit problems.
An approximate solution to a very similar problem is proposed here:
http://www.mit.edu/~hauser/Papers/Hauser_Urban_Liberali_Braun_Website_Morphing_May_2008.pdf http://www.mit.edu/~hauser/Papers/Hauser_Urban_Liberali_Brau...
Once you see the logic of this approach, it's really shocking that A/B testing companies have not implemented it. It's really the only way to think about optimal decision making under uncertainty.