7 ms·
The purpose of an A/B test isn't to always show the best performing result, it's to perform a _controlled scientific experiment_ with a control group, from whic
by err4nt 7y ago
The purpose of an A/B test isn't to always show the best performing result, it's to perform a _controlled scientific experiment_ with a control group, from which you can learn things.
Also, I work in this field and I will just say that people _do_ behave differently based on traffic source: i.e. users coming from Facebook behave alike, but different than traffic from Reddit who act similarly to each other. If you were running a self-optimizing thing like this it _would_ make sense to split it up by the different traffic sources and handle them separately.
- elehack 7y agoYes. Bandits will often converge more quickly to the optimal strategy, but it is much more difficult to understand why that strategy is optimal and generalize from the bandit outcomes to predict future performance and performance of other strategies. It isn't impossible - bandits are seeing adoption in medical trials to avoid precisely the problem discussed - but the standard experiment design and analysis techniques you learn in a decent college statistics class or introductory statistics text no longer apply. That's one of the beauties of A/B testing: while it does require substantial thought to do well, the basic statistics of the setup are very well-understood at this point.
- orasis 7y agoI disagree. I’ve spent a lot of time staring at bandit outcomes and usually they match some sort of intuition of why a variant might be exceptional.
- comicjk 7y agoThat could be post-hoc reasoning, though. It would be interesting to pre-register your hypotheses, or see whether you could tell bandit outcomes from random ones.
- edmundsauto 7y agoIsn't this problem also an issue when people talk about transferring what they learn from one test to another test? That is frequently cited as a benefit of A/B testing.
- orasis 7y agoSure it’s post-hoc reasoning, but it doesn’t matter because I’m not trying to invalidate a hypothesis. I’m looking for variants that win. When I find one that wins I look at it and try to add more of the same flavor to the product. This process works.
- taeric 7y agoThis is literally the logical fallacy. You could get lucky. Maybe you have obvious gains to chase. But bad logical arguments are bad because they never work forever. They are corrupted heuristics that can get you in trouble without critical thinking. Edit: added in forever. Phone dropped some wording I originally had. I think.
- orasis 7y agoCall it a genetic algorithm if you like. I’m looking for incremental wins in a world of infinite possibilities, not truth.
- taeric 7y agoIncremental wins can still lead to dead ends. My phrasing was off in my post. I meant to say that the fallacies aren't that the tactics never work. Just that they can stop working without you really realizing it. A heuristics that can lead you down a dead end. By all means, keep doing it if it is working for you. But don't confuse it as good advice. And stay vigilant.
- orasis 7y agoProducts exist in human reality not some science paper. There are no absolute truths, everything dead-ends eventually. It’s like trying to prove that one set of genes is better than another for future survival - an impossible task.
- edmundsauto 7y agoBut for results to generalize or to understand why, the confounders must be accounted for in the randomization. This is really hard to do well -- there are often subtle influences that aren't sufficiently understood how they impact these non-linear systems. What makes someone convert? A million different factors; changing the color of a button in one context doesn't necessarily tell me much about how people would respond to that experience in another context. It's easy to underestimate how complex things are, because we only see some superficial aspects of e.g. a user/software interaction model. This flaw is down to how our brains work -- ref "What you see is all there is".
- citrablue 7y agoWhat you said is correct, but I'd like to point out that controlled scientific experiments may not be the right approach for e.g. optimizing conversions on a website. The reason is that websites are a dynamic environment. All things equal, better controlled experiments are great. However, visitor behavior, especially when from dynamic sources (google serps change weekly), changes all the time. And that's why I prefer MAB over A/B tests -- A/B tests don't adapt to a dynamic system, so we often wish away the changes in the system to use it. Does anyone go back and re-test their biggest wins?
- err4nt 7y ago> Does anyone go back and re-test their biggest wins? Yes, absolutely! We do research first, then come up with simple, well-controlled tests. Once we have a winner we can either lock it in, but often we continue to research and experiment on the new knowledge we gained. A hefty minority of the tests I implement build on past wins to further flesh out what works and what doesn't with knowledge and the data to back it up.
- JohannesH 7y agoVery interesting. It seems to me that doing incremental work like this might end up in a local minima/maxima. Do you have any advice on how to avoid pitfalls like that? Are you testing radically different ideas along with your incremental improvements?
- orasis 7y agoFrom a workflow perspective, MAB is a bit difficult to find radical improvements from. The radical improvements come from fundamental design changes, of which it would be very expensive to create a bunch of radically different variants. MAB is best used where you can generate a bunch of variants cheaply and hope for a 30% gain.
- JohannesH 7y agoCool, thanks.
- arendtio 7y ago> The purpose of an A/B test isn't to always show the best performing result The key part is 'not always'. Using A/B tests just to find the better performing version is a valid use case too. And if you have the traffic, multi-armed bandits are a nice way of automating the whole procedure. Even scientifically speaking there is nothing wrong with them. Their biggest issue is that they require a lot of traffic for significant results.
- joshuamorton 7y agoIn practice, MAB should require less traffic than an A/B test for the same level of significance. Although it's much more difficult to rigorously describe that significance level with a MAB.
- antidesitter 7y agoThat’s just a contextual multi-armed bandit.
- throwawaymath 7y ago> Also, I work in this field and I will just say that people _do_ behave differently based on traffic source: i.e. users coming from Facebook behave alike, but different than traffic from Reddit who act similarly to each other. Out of curiosity, what is your hypothesis for explaining this difference in behavior? Would you say it's primarily due to differing contexts in which a link is posted, or differing populations on each platform? Or maybe these both contribute about equally? Stated another way: would you expect the same individual to behave differently coming from Facebook versus coming from reddit, if they happen to be a user of both?
- orasis 7y agoDifferent demographics, different intent, and different mental context all play a factor.
- parksy 7y agoNot a statistician by any means, but could traffic source be a factor that's evaluated alongside conversion by a bandit algorithm when calculating the chance to show a particular option? Or other factors as well (detected device capabilities, users location, etc?) These could just be weighting factors so instead of a single % chance per option, every time there's a successful interaction the victory is spread across the factors for that option. New users could be shown the option where the chance for each option is weighted by the factors they match. Is this a valid approach or would it introduce some kind of selection bias?
- orasis 7y agoYes, of course. You can use a contextual MAB and include the traffic source in the context. Contextual MABs are much more complicated and expensive to implement though.
- paulddraper 7y ago> The purpose of an A/B test isn't to always show the best performing result, it's to perform a _controlled scientific experiment_ with a control group, from which you can learn things. Occasionally there is pure scientific interest. But far more frequently, the purpose of the A/B is to optimize the outcome. This is why Google Analytics has exclusively chosen multi-armed bandit for its its A/B test framework.
- mobjack 7y agoIt doesn't have to be a pure science interest. If you want to use the results from one test to inform what to test next, then A/B tests are better optimizing for truth. Most website changes dont make a significant difference in conversion. If you use MAB, then you dont know if the winner is really better or the result of random variation.
- paulddraper 7y ago> If you use MAB, then you dont know if the winner is really better or the result of random variation. No, you know the outcome of MAB is the best. If it's clearly the best, it converges quickly. If it's less clearly the best, it converges slowly. Either way though, you haven't lost any more conversions that necessary to find that out, or needed to put in non-mathematically fail safes.
- bduerst 7y agoWouldn't a simple solution be to create a handler that runs a MAB for each traffic source? If you're worried about the site changing for people between visits from different traffic sources, you can cookie their MAB/test-assignment on the first visit.
- err4nt 7y agoIf your ultimate goal is just to convert, convert, convert, and you can generate enough different content in your content strategy to sell different types of people (usually this content strategy is what's missing) you can set up tracking so it tracks everything you can imagine (did the user scroll down and then scroll all the way back up to the top? How long was the page loaded before the user scrolled the first time? and all kinds of stuff is trackable) you can split the traffic by their origin (because traffic sources seem to behave alike) and I know a guy who does Lead Qualification and his setup he can predict with a 90% accuracy whether a user will convert or not after 1000 visitors from that traffic source. And he's tracking about 30 variables to come to that prediction. These things are easy to track, but if you can track it but your content strategy has no alternate content or anything to do, having that ability is useless. Some things that you could do to alter a sales funnel to try to convert a customer: - if you think they need more priming before they're ready to buy, add more pages into the funnel for that user - if you think they're in a place to buy immediately, make sure you show them a Call to Action right then so they can buy - swap entire blocks of text based on what kind of user we think you are - change the length of pages, (add or rearranging elements on the page to highlight different things to different users based on what we think they care about more) - etc You can also often just ask users to self-identify a 'role' that is useful for content strategy, like "Are you a [student] or [teacher]" and people will click the button, and even expect the page to change and customize to that experience. All this stuff is so easy to do technologically, but you have to have a content strategy that actually makes use of the insights available.