6 ms·
If the p-values were accurate and averaged around 0.05, ~95% of results should be reproducible. That only 36% were points to deep, fundamental errors.
by jsprogrammer 11y ago
If the p-values were accurate and averaged around 0.05, ~95% of results should be reproducible.
That only 36% were points to deep, fundamental errors.
- gwern 11y agoNo. P-values don't work that way and don't mean what you think they mean. Read OP or heck, any of the classics like "Why most published research findings are false" http://dx.plos.org/10.1371/journal.pmed.0020124 http://dx.plos.org/10.1371/journal.pmed.0020124 (36% may or may not be bad, but you can't know without additional stuff like power or prior probability of hypotheses being true; p-values have no intuitive meaning and aren't an answer to any question that people are asking, which is a major reason why Bayesian approaches can be useful. And from a Bayesian perspective, I find 36% totally unsurprising - if anything, substantially better than I had expected given the gross underpowering of most psych studies, the statistical-significance publication filter, and the dubiousness of most hypotheses.)
- jsprogrammer 11y agoA proper rebuttal would show what a p-value actually is and how it differs from what I claimed. Now, since a p-value is exactly what I previously claimed, you obviously can't do that. I'm not even sure what you are arguing against me here.
- Dylan16807 11y agoThe p-value is the chance of a false positive. But you don't know what the rate of true positives is, or the rate of false negatives. In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. But the reproduction rate when there actually is an effect is not 95%. Depending on sample size, I might get a true positive 20% of the time and a false negative 80% of the time, or I might get a true positive 99.8% of the time and a false negative .2% of the time. So the average reproduction rate, where an effect actually exists, can be almost any number between 5 and 100. There is no reason to assume it will be 95%. So the average reproduction rate, where some effects are real and some are imaginary, will almost certainly not be exactly 95%, and that is not a problem in and of itself. (And when you talk about an average p-value of .05, that sounds like only publishing positive results, which is blatantly going to fail reproduction. 100 false hypotheses -> 5 publications, all false positives -> 5% reproduction rate)
- jsprogrammer 11y ago>In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%. This is the world p-value assumes and is therefore the only one worth considering in relation to my comment. If an experiment is not well-formed then of course you won't see reproduction at the expected rate. This is what I'm referring to when I say that the low reproduction rate points to deep, fundamental flaws in the experiments. I agree that the reproduction rate will never be exactly 95% (or 1 - p) due to the discrete nature of experimentation [that's why I used a ~ in front :)], but the reproduction rate of a well-formed experiment should very closely track 1 - p.
- Dylan16807 11y ago>This is the world p-value assumes and is therefore the only one worth considering in relation to my comment. I'm not sure if that was clear enough. In that world, no one has ever had a hypothesis that was correct. The whole field is useless, measuring things that are wrong and getting the occasional false positive. You can talk about that world if you want, but it has no connection to reality. It's not p-values that assume that world, it's your misunderstanding of p-values. >If an experiment is not well-formed then of course you won't see reproduction at the expected rate. This is what I'm referring to when I say that the low reproduction rate points to deep, fundamental flaws in the experiments. Experiments don't have to have enormous sample sizes to be well-formed. That's the whole point of having a cutoff value. It's not like an experiment that reproduces 80% of the time disproves the result the rest of the time, it just doesn't quite reach .05 on those trials >the reproduction rate of a well-formed experiment should very closely track 1 - p I'm suspicious of this. I don't have time to do the math right now, but an experiment that averages .01 might clear a .05 hurdle far more than 99% of the time, and would definitely be well-formed. And if you set a hurdle at .01 it would only clear it half the time, but it would still be well-formed.
- jsprogrammer 11y agoHypotheses can never be proven to be correct. I don't want to be in any world where it is believed that a hypothesis is or could be correct. This is a fundamental tenant of science. All that can be done is to reject hypotheses. You (along with Gwern) have now claimed that I don't understand p-values, but you present no alternative understanding. The reason, of course, is that when you look at the mathematics behind p-value, it is obvious that it is exactly as I claim. Edit to address your edit: >I'm suspicious of this. I don't have time to do the math right now, but an experiment that averages .01 might clear a .05 hurdle far more than 99% of the time, and would definitely be well-formed. And if you set a hurdle at .01 it would only clear it half the time, but it would still be well-formed. You are right that you need to be careful here about what you are comparing across instances. There will be variability since you are only sampling a distribution (most likely at a very low rate) and not observing the entire distribution (which, for continuous distributions, is impossible).
- washedup 11y agoI agree, 36% ain't too bad. But,it requires that any literature you use in your research should have been reproduced a few times by other researchers.
- kgwgk 11y agoLet me join the club of people claiming that you don't understand p-values. It's not clear if you are talking about the rate of reproduction for a subset of the possible experimental outcomes (those rejecting the null at the alpha=5% level) or for the whole set. When the null hypothesis is true (remember that there are fields where this is the norm, v.g. ESP), you would only reproduce (reject again) 5% of the rejections. Of course you would reproduce (non-reject for the second time) 95% of the non-rejections. The global reproduction rate would be 0.95 x 0.95+0.05 x 0.5=0.905 (90.5% doesn't look ~95% either). When the null hypothesis is not true, the probability of reproducing (in either sense) the result of a test depends on the effect size. If the effect is huge, the test will be rejected with probability ~100% and the result will be reproduced with probability ~100%. Or maybe you mean by reproducing "getting a lower p-value" in the the second trial? If the null is true, the probability of getting p-value2<p-value1 is precisely p-value1. If the null is not true, it will depend on the effect size. If you assume the effect size is the observed one, you expect p-value2 to be smaller than p-value1 with probability 50%.
- jsprogrammer 11y ago>If the null is true, the probability of getting p-value2<p-value1 is precisely p-value1. Precisely. P-value is only defined when the null hypothesis is true. Apparently everyone else is overlooking this fact.
- kgwgk 11y agoI don't see how does it contradict anything that I (and everyone else) wrote. You seem to agree that when the null is true and the original result was p-value=0.05, the probability of reproducing the result (getting p-value<0.05 on a second trial) is 5%. This seems incompatible with your original claim: "If the p-values were accurate and averaged around 0.05, ~95% of results should be reproducible." Could you explain exactly what do the following mean: p-values were accurate (that the null is true?) p-values averaged around 0.05 (that you're taking the subset of outcomes with p-value 0.05?) ~95% of results should be reproducible (that if you take the previous subset you will get p-value<0.05 always in 95% of them? or exactly 95% of the time in all of them?)