3 ms·
This post is overconfidently written, and its thesis is wrong. A few others have mentioned CUPED (Kohavi et al. 2013), a way of doing exactly what this blog po
by goodside 5y ago
This post is overconfidently written, and its thesis is wrong.
A few others have mentioned CUPED (Kohavi et al. 2013), a way of doing exactly what this blog post says you can’t. CUPED is a great starting point for anyone new to the subject of pre-experiment control variates.
The post is bold enough to describe any use of control variates is a “fallacy” but somehow doesn’t mention the most famous method for doing this (CUPED). The author is only interested in linking to their own prior blog posts, and does not engage with any reasonable (or attributed) argument for the method they’re dismissing.
The post’s main argument is that using CUPED-style control variates provides no benefit when N is sufficiently large, so we should just make N bigger — as though “sufficient N” grew on trees.
The next argument is less silly, but still wrong: They note observed differences in pre-treatment effects are “random fluctuations”, and claim they thus cannot be used for anything. Their argument ignores the reason pre-treatment measurements are subtracted in the first place: In CUPED, we acknowledge pre-treatment observations are undesirable noise, but presume this noise already contaminates our post-treatment measurements. We collect pre-treatment measurements exactly because we want them gone — the pre-treatment measurements tell us how much meaningless noise to subtract.
As a reductio ad absurdum, imagine an “A/B test” (an RCT) of a drug that ostensibly makes humans taller. Would it seem reasonable to you, as a participant, if the doctors never ask what your initial height is, and only measure you once the trial is done?
- Maro 5y agoThanks for the pointer to CUPED! I wasn't aware of it. After a quick scan, it seems to me that CUPED is supposed to be run at a per-unit level (ie. per user normalization), to reduce variance, which seems to be a bit different than the fallacy I describe here (computing lifts from the T and C group's "before" and "after" overall mean separately and subtracting the lifts, which is what I observed and triggered me to write this post). (I wrote the post.)
- johnmyleswhite 5y agoIf you work through the math for CUPED, you'll see that only the aggregate correlation term can be interpreted as being "per-unit" -- all of the other terms in the equation are the same terms you use in your post. Insofar as I think your intuition is leading you somewhere, I think it's leading you towards a realization that a "diff in diff" approach rather than regression adjustment can increase variance in some settings. But regression adjustment is provably better in essentially all circumstances: the only settings in which it is ever worse than no adjustment are outlined clearly in https://projecteuclid.org/journals/annals-of-applied-statistics/volume-7/issue-1/Agnostic-notes-on-regression-adjustments-to-experimental-data--Reexamining/10.1214/12-AOAS583.full https://projecteuclid.org/journals/annals-of-applied-statist...
- Maro 5y agoScanning various CUPED related pages, I read that it's a way to reduce the variance, and hence p-value. But CUPED is not changing the lift value (difference [or ratio] in means) between T and C itself (or at least, not in the examples I see). The fallacy I describe computes different means from historic lifts and substracts those. Ie. on the third table, the lift is 4.7%, not -1.6%.