6 ms·
> I wanted Optimizely to say it was 100% significant That's not how statistical significance works...
by adevine 11y ago
> I wanted Optimizely to say it was 100% significant
That's not how statistical significance works...
- mangeletti 11y agoYou should contact Optimizely and let them know right now. https://www.optimizely.com/contact/ https://www.optimizely.com/contact/ They'll need to reeducate their statisticians right away!
- tamana 11y agoOptimizely did an A/B test that showed that customers respond better to rounded-up numbers. ;-) the real world is messy. Years ago my team's statistician did a competitive review of various AB test apps, and reported various ways in which the UIs make statistically invalid statements to the user.
- xixi77 11y agoThey probably do, but why would they want to do so, when doing it would make their results appear less conclusive?
- norkakn 11y agoThe basic model that Optimizely uses is a Z-Test approximation of a binomial distribution. To run a proper experiment with that model, you should be calculating the sample size ahead of time, and then run it. Each visitor should be independent, and not affected by things like the day of the week, or the time of it. The end result tells you if the distributions are different, but not as much as one would think about the size of the differences. It also can't be 100. The normal distribution has an infinite range, so a finite limit can never capture 100% of it. Optimizely is in a rough spot. People don't like having to think through experimental design, and they are really, really bad at reasoning about p-values. To try to fix the people part, they came out with the sequential stopping rule stuff (their "stats engine), but they never really published much justifying it. The other alternative would be to move the experiments into a Bayesian framework, but that has a lot of it's own problems. When they acquired Synference, that was one of the likely directions to take (along with offering bandits), but that didn't work out and those guys have since left.
- gingerlime 11y ago> but they never really published much justifying it Not that I'm trying to defend Optimizely (I'm not a huge fan, but for other reasons...). I can't vouch for the quality either, but they did publish something about it[0] - that at least looks quite scientific. Happy to read any critique of course. [0] http://pages.optimizely.com/rs/optimizely/images/stats_engine_technical_paper.pdf http://pages.optimizely.com/rs/optimizely/images/stats_engin...
- norkakn 11y agoLatex is a wonderful way to make a marketing paper look like a scientific one. It doesn't accurately describe the method, but that isn't really its purpose. It's a more technical description of the blog post, meant for people using the product to understand some of the tradeoffs and get more accurate results. They are still having people make very fundamentally flawed assumptions about the data, which results in incorrect conclusions, and they are still not presenting the results in a way that people correctly interpret them. That being said, those are really hard to solve, and models that would try to correct for them would likely require a lot more data and be overly conservative for more people. What are your reasons for disliking Optimizely?
- yummyfajitas 11y agoIt may be marketing but I was able to implement a sequential A/B test based on it. Admittedly, I did need to do some work beyond merely copy/pasting an algorithm, but all I really needed to do was read their paper and some citations. I do believe that this document does describe a viable frequentist test and my implementation of it worked pretty well. Disclaimer: I do stats work at VWO, an Optimizely competitor. (Also if you want to read our tech paper, here it is: https://cdn2.hubspot.net/hubfs/310840/VWO_SmartStats_technical_whitepaper.pdf https://cdn2.hubspot.net/hubfs/310840/VWO_SmartStats_technic... This describes our Bayesian approach, which we believe to be less likely to be wrongly interpreted by non-statisticians.)
- leo_pekelis 11y agoHi, this is Leo, Optimizely's statistician. If you're looking for a more scientific paper, maybe take a look at this one we wrote recently: http://arxiv.org/abs/1512.04922 http://arxiv.org/abs/1512.04922 Should have everything you would ever want to know about the method. I agree with you that the problem of inference and interpretation between A/B data, algorithms, and the people who make decisions from them is a hard one and worth working on. That said, I do think the two sources of error our stats engine addresses - repeatedly checking results, and cherry picking from many metrics and variations - did make progress in having folks correctly interpret A/B Tests. This did result in more conservative results, but the benefit was that the variations that do become significant are more trustworthy. I think this was absolutely the right tradeoff to make for our customers, and trustworthyness is a pretty important aspiration for stats/ML/data science in general. Of course I did write the thing, so I'm not very impartial.
- Quinner 11y agoThere's a reason you'll only see their number say >99%
- gk1 11y agoYou're wrong. What you are thinking about is Optimizely's "Chance to Beat Baseline" number. That's different from the statistical significance, which is a setting you can change on the Settings page. Being smug and condescending really backfires when you don't know what you're talking about.
- mangeletti 11y agoBro... http://cl.ly/1b0B3Y3o1w09 http://cl.ly/1b0B3Y3o1w09 > Being smug and condescending really backfires when you don't know what you're talking about. How's that working out for you?