4 ms·
Can you help me understand why we would use sample unit engineering/bootstrapping? Imagine if we don't care about between subjects variance (and thus P-values i
by authorfly 2y ago
Can you help me understand why we would use sample unit engineering/bootstrapping? Imagine if we don't care about between subjects variance (and thus P-values in T-tests/AB tests), in that case, it doesn't help us right...
I just feel intuitively that it's masking the variance by converting it into within-subjects variance arbitrarily.
Here's my layman-ish interpretation:
P-values are easier to obtain when the variance is reduced. But we established P-values and the 0.05 threshold before these techniques. With the new techniques reducing SD, which P-values directly interpret, you need to counteract the reduction in SD of the samples with a harsher P-value in order to obtain the same number of true positive experiments as when P-values were originally proposed. In other words, allowing more experiments to have less variance in group tests and result in more statistical significant if there is an effect size is not necessarily advantageous. Especially if we consider the purpose of statistics and AB testing to be rejecting the null hypothesis, rather than showing significant effect sizes.
- kqr 2y agoLet's use the classic example of "Lady tasting tea". Someone claims to be able to tell, by taste alone, if milk was added before or after boiling water. We can imagine two versions of this test. In both, we serve 12 cups of tea, six of which have had milk added first. In one of the experiments, we keep everything else the same: same quantities of milk and tea, same steeping time, same type of tea, same source of water, etc. In the other experiment, we randomly vary quantities of milk and tea, steeping time, type of tea etc. Both of these experiments are valid, both have the same 5 % risk of false positives (given by the null hypothesis that any judgment by the Lady is a coinflip). But you can probably intuit that in one of the experiments, the Lady has a greater chance of proving her acumen, because there are fewer distractions. Maybe she is able to discern milk-first-or-last by taste, but this gets muddled up by all the variations in the second experiment. In other words, the cleaner experiment is more sensitive, but it is not at a greater risk of false positives. The same can be said of sample unit engineering: it makes experiments more sensitive (i.e. we can detect a finer signal for the same cost) without increasing the risk of false positives (which is fixed by the type of test we run.) ---- Sometimes we only care about detecting a large effect, and a small effect is clinically insignificant. Maybe we are only impressed by the Lady if she can discern despite distractions of many variations. Then removing distractions is a mistake. But traditional hypothesis tests of that kind are designed from the perspective of "any signal, however small, is meaningful." (I think this is even a requirement for using frequentist methods. They neef an exact null hypothesis to compute probabilities from.)
- authorfly 2y agoThank you. I'll have to think about it a bit more but I appreciate you response