5 ms·
One of the most frustrating results I found is that A/B split tests often resolved into a winner within the sample size range we set; however if I left the spli
by sunir 2y ago
One of the most frustrating results I found is that A/B split tests often resolved into a winner within the sample size range we set; however if I left the split running over a longer period of time (eg a year) the difference would wash out.
I had retargeting in a 24 month split by accident and found it didn’t matter after all the cost in the long term. We could bend the conversion curve but not change the people who would convert.
And yes we did capture more revenue in the short term but over the long term the cost of the ads netted it all to zero or less than zero. And yes we turned off retreating after conversion. The result was customers who weren’t retargeted eventually bought anyway.
Has anyone else experienced the same?
- bdjsiqoocwk 2y ago> One of the most frustrating results I found is that A/B split tests often resolved into a winner within the sample size range we set; however if I left the split running over a longer period of time (eg a year) the difference would wash out. Doesn't that just mean there's no difference? Why is that frustrating? Does the frustration come from the expectation that any little variable might make a difference? Should I use red buttons or blue buttons? Maybe if the product is shit, the color of the buttons doesn't matter.
- admax88qqq 2y ago> Maybe if the product is shit, the color of the buttons doesn't matter. This should really be on a poster in many offices.
- cwillu 2y ago“It looks awful, and it works” (apologies to Buckley's)
- sunir 2y agoMy frustration is the a/b split tests never seemed to net to anything in the long term even after confidence was reached. It made me question the entire process; but I understand the math so it’s confusing to me.
- kqr 2y ago> We could bend the conversion curve but not change the people who would convert. I think this is very common. I talked to salespeople who claimed that customers on 2.0 are happier than those on 1.0, which they had determined by measuring satisfaction in the two groups and got a statistically significant result. What they didn't realise was that almost all of the customers on 2.0 had been those that willingly upgraded from 1.0. What sort of customer willingly upgrades? The most satisfied ones. Again: they bent the curve, didn't change the people. I'm sure this type of confounding-by-self-selection is incredibly common.
- Adverblessly 2y agoObviously it depends on the exact test you are running, but a factor that is frequently ignored in A/B testing is that often one arm of the experiment is the existing state vs. another arm that is some novel state, and such novelty can itself have an effect. E.g. it doesn't really matter if this widget is blue or green, but changing it from one color to the other temporarily increases user attention to it, until they are again used to the new color. Users don't actually prefer your new flow for X over the old one, but because it is new they are trying it out, etc.
- sunir 2y agoMaybe. Retargeting is unlikely to create novelty.