2 ms·
What I believe the post is recommending is that you attempt to characterize your population, and then select your groupings based on that characterization. The
by icegreentea 11y ago
What I believe the post is recommending is that you attempt to characterize your population, and then select your groupings based on that characterization. The goal is that the distribution of your two groups across all possible metrics you can measure are as close to equal as possible. Some trivial (and obviously extreme) examples of the medical field could be:
a) If you're doing a study with 50/50 males and female participants, you'd probably reroll your control/active group distributions if it was far off from 50/50. Hell, you'd probably just split males and females separately.
b) If you had a twin study, the only way to divide up your population that makes sense is by splitting up your twins. Nothing else makes sense.
c) Imagine you wanted to test how Adderall effects the test studying/taking abilities of a variety of college aged students. You would do your best to make sure that your control/active populations have similar distributions in your test metric (or test metric proxy such as IQ) prior to starting test.
Example c) is the closest to the case that the post is talking about. When you're trying your absolute best to maximize the power of your test, usually want to take into account what you know, or what you assume about the thing your testing. In the case of website A/B testing, you're making an assumption (that is possibly unfounded) that reaction to A is an okay proxy for reaction to B.
This is an assumption - nearly all statistics is based on assumptions. Nearly all exercises in powering and designing experiments is based on assumptions and iteration.