4 ms·
10% of the time, we choose a lever at random. The other 90% of the time, we choose the lever that has the highest expectation of rewards. There is a p
by 3dfan 11y ago
10% of the time, we choose a lever at random. The
other 90% of the time, we choose the lever that has
the highest expectation of rewards.
There is a problem with strategies that change the distribution over time: Other factors change over time too.
For example let's say over time the percentage of your traffic that comes from search engines increases. And this traffic converts better then your other traffic. And let's say at the same time, your yellow button gets picked more often then your green button.
This will make it look like the yellow button performs better then it actually does. Because it got more views during a time where the traffic was better.
This can drive your website in the wrong direction. If the yellow button performs better at first just by chance then it will be displayed more and more. If at the same time the quality of your traffic improves, that makes it look like the yellow button is better. While in reality it might be worse.
In the end, the results of these kinds of adaptive strategies are almost impossible to interpret.
- zeckalpha 11y agoIt's possible to weight conversions by frecency to account for this, rather than using frequency alone.
- kevin_nisbet 11y agoI don't know if this is the case if I understand this algorithm correctly. Say Yellow is 50%, and Green is a 65% success rate after the behavior change, but green is 30% before the behavior change. By sending 90% of traffic towards yellow, it's ratio will normalize towards the 50% once it has enough traffic. By sending 10% of traffic randomly, eventually the green option will reach 51%, and start taking a majority of traffic, which then will cause it to normalize at it's 65%, and be shown to a majority of users. I think the problem might be, if you run this with a sufficiently high volume or for a long period of time, that if a behaviour change takes place it will take a long time to learn the new behaviour. Or if two options aren't actually different, it may continually flip back and fourth between two options. Also, to me, the concept of A/B testing certain things may also have an undesired consequence. For example, I order from amazon every day, but today the but button is blue, what does that actually mean? And I go back to the site later and it's yellow again. There are still many people who get confused by seemingly innocuous changes with the way their computer interacts with them.
- TuringTest 11y ago> Also, to me, the concept of A/B testing certain things may also have an undesired consequence. For example, I order from amazon every day, but today the but button is blue, what does that actually mean? And I go back to the site later and it's yellow again. There are still many people who get confused by seemingly innocuous changes with the way their computer interacts with them. Proper A/B tests are supposed to be done on a per-unique-user basis. If you access the shop from the same device or user account, a well-done A/B test should consistently show you the same interface.
- snowwrestler 11y agoUntil that particular test ends, and another one begins.
- kevin_nisbet 11y agoI agree. Sorry in this context I was referring specifically to the "in 20 lines of code article" which I don't believe had this control to it.
- metafunctor 11y agoThis is why you should segment traffic and run separate tests for each segment, whether you're using an A/B testing or a multi-armed banding algorithm.
- zardeh 11y agoIndeed, otherwise you might be guilty of have Simpson's paradox either way.