2 ms·
This is a good overview of the multi-arm bandit problem [1], but the author is far too dismissive of A/B Testing. First of all, the suggested approach isn't al
by ryporter 11y ago
This is a good overview of the multi-arm bandit problem [1], but the author is far too dismissive of A/B Testing.
First of all, the suggested approach isn't always practical. Imagine that you are testing an overhaul of your website. Do you want daily individual visitors to keep flipping back and forth as the probabilities change? I'm not sure if the author is really suggesting his approach would be a better way to run drug trials, but that's clearly ridiculous. You have to recruit a set of people to participate in the study, and then you obviously can't change what drug you're giving them during the course of experiment!
Second, it ignores the time value of completing an experiment earlier. In the exploration/exploitation tradeoff, sometimes short-term exploitation isn't nearly as valuable as wrapping up an experiment so that your team can move to new experiments (e.g., shuting down the old website in the previous example). If a company expects to have a long lifetime, then, over the a time frame measured in weeks, exploration will likely be relatively far more valuable.
[1] https://en.wikipedia.org/wiki/Multi-armed_bandit https://en.wikipedia.org/wiki/Multi-armed_bandit
- davkap92 11y agoRegarding your first point, not sure if author covered it, but in Google Content Experiments with multi bandit approach cookies are stored, so user who sees variation b will keep seeing b while the experiment is running
- mplewis 11y agoThis is the approach every good AB testing service/framework uses.