4 ms·
I assume you mean Null Hypothesis Significance Testing. Can you elaborate or provide sources as to why this destroys adopters?
by rmellow 8y ago
I assume you mean Null Hypothesis Significance Testing. Can you elaborate or provide sources as to why this destroys adopters?
- notafraudster 8y agoTo be brief: The problems associated with NHST are well documented (three that are especially awful: multiple comparisons lead to uncontrolled rate of false positives; large sample sizes leading to statistically significant but substantively irrelevant effects being "accepted"; hard thresholds of statistical significance are arbitrary and the difference between a significant effect and insignificant effect is rarely significant). In addition, there are problems that emerge from using NHST as part of an A/B test: namely, using live, rolling samples to identify significant results rather than prespecifying sample sizes; and making atomic changes as though they are additive without re-testing joint hypotheses of multiple changes. But when you come down to it, the question you want to ask is "Is B better than A?" and the question you do ask is "If A were truly no better than B, how often would a sample of size n drawn using the same sampling procedure I think I'm using produce the impression than B is as much better than A as it is apparently in my observed data?", and two problems are that these aren't the same question and almost no one knows they're doing the latter. To be totally fair, one of the most common problems with NHST (the null hypothesis is patently absurd) isn't necessarily a problem in the A/B UX case. Not sure which of these in specific the grandparent is referring to, but I suspect they and I are on the same page in general.
- nonbel 8y agoYou may be interested in distinguishing between enumerative and analytic studies as described here: https://s3.amazonaws.com/wedi/www/Articles/b21d561f-9aea-4b8d-8757-d70964ae13b5.pdf https://s3.amazonaws.com/wedi/www/Articles/b21d561f-9aea-4b8...
- nonbel 8y agoThere is always a difference between condition A and condition B. It will be detected with careful enough measurement and/or large enough sample size (spend enough money). Unless you can predict what this difference is expected to be with some precision beforehand it is not possible to legitimately attribute it to your favorite reason. I would start here: https://www.jstor.org/stable/186099 https://www.jstor.org/stable/186099