3 ms·
A big problem is that there's often no objective way to pick the "right" statistical test. This gives experimenter the freedom to choose the statistical procedu
by yaroslavvb 17y ago
A big problem is that there's often no objective way to pick the "right" statistical test. This gives experimenter the freedom to choose the statistical procedure that favors positive conclusion. Here's a classical example when the statistician has the freedom in choosing between binomial and negative binomial test
Two experimenters contrast treatments A and B. They both have A preferred to B in first 5 patients, and B preferred in the 6th. First experimenter planned to run 6 experiments and count the number of successes, so they get P-value 0.11 for the hypothesis that A is better. The second experimenter planned to run comparisons until B is preferred, up to 6, and got P-value of 0.03.
(see Appendix A of http://www.annals.org/content/130/12/995.full.pdf+html http://www.annals.org/content/130/12/995.full.pdf+html)
A realistic example of this issue coming up
http://www.jstor.org/pss/2336980 http://www.jstor.org/pss/2336980
Another example is choosing between one-tailed and two-tailed t-test. When you ask for the probability of effect being as extreme as observed x under null hypothesis, should you ask for probability of effect>x or |effect|>|x|?
Eliezer Yudkowsky goes into some discussion on this
http://lesswrong.com/lw/1gc/frequentist_statistics_are_frequently_subjective/ http://lesswrong.com/lw/1gc/frequentist_statistics_are_frequ...
The most illustrative example of subjectivity of hypothesis testing is probably the issue of testing strings for randomness. There are many tests for testing whether a particular string of bits is generated by a Bernoulli process, with not one having a legitimate claim to being "the right one"
One way to remove the bias is three-way triple-blind testing, ie to measure effects of existing, new, and placebo treatments, have statistician analyze datasets while blinded to their true labels.