4 ms·
Even if you assume best intentions, A/B testing often comes down to a variation of a t-test, the interpretation of which can be fuzzy too - https://www.amstat.o
by 0x31a 7y ago
Even if you assume best intentions, A/B testing often comes down to a variation of a t-test, the interpretation of which can be fuzzy too - https://www.amstat.org/asa/files/pdfs/P-ValueStatement.pdf https://www.amstat.org/asa/files/pdfs/P-ValueStatement.pdf (TLDR - when it comes to statistical significance you can be either wrong or vague).
- srean 7y agoWhy t-test in particular ? That would be applicable as long as one has two Gaussian populations with the same variance and potentially different means. In my experience, those preconditions are violated more frequently than they are met, unless some extra steps are taken to make the data Gaussian with equal variance. Central limit theorem helps sometimes if the true variance of the populations are well behaved (not very large).
- 0x31a 7y agoYou don't need equal variances - https://en.wikipedia.org/wiki/Welch%27s_t-test https://en.wikipedia.org/wiki/Welch%27s_t-test
- mattkrause 7y agoIt's also surprisingly robust to non-Gaussianity.
- srean 7y agoNot if it has a heavier tail, which is common in practice. It is disastrously brittle in those cases, mainly because sample variance is a lot of garbage in those cases.
- mattkrause 7y agoEhh, it's more complicated than that. Asymmetrical distributions don't fare well, in part because the mean probably isn't the thing you want to be comparing anyway. Mixture distributions or data contaminated with outliers are also trouble. On the other hand, a bunch of simulation studies have found that t-tests work pretty well for symmetric distributions, even those with moderately heavy tails like a t(6) distribution. Moreover, the power declines a lot faster than the Type I error rate (which, for whatever reason, people seem to worry about more). I've gotten a few paper reviews where someone has made the "but it's not normal" argument. Redoing the analysis with permutation or randomization tests (or both, for one stubborn reviewer) almost never changes the resulting p-values by more than a few percent. I don't know if this reflects the fact that they're usually quite similar or that we're just fairly judicious in how we analyze data.
- srean 7y agoI would agree that its a lot more resistant than one would think, but even an innocuous Gaussian with a heavy tailed contamination confuses it. For cases where one needs too test "is this bigger than that ?" I test for stochastic dominance rather than equality of mean.
- srean 7y agoThat would be Welch test and not t-test