6 ms·
From my experience, scientists, -at least in biology, where like in sociology you might have a lot of noise to deal with-, have an internal intuition that a sin
From my experience, scientists, -at least in biology, where like in sociology you might have a lot of noise to deal with-, have an internal intuition that a single paper with a significant result does not mean that we have found the truth. The recent study which reported a reproducibility in sociology of about 36% strikes me as pretty accurate.
I think the scientific system can work with that. It means that if you build follow-up experiments based on a single paper there is a good chance that the experiment fails. In some way, the scientific system of publishing is self-correcting in this regard, because you can then cast doubt on the previous paper, which is easier to publish than if you only have a fresh negative result (p-value > threshold).
There is no way to know how many people tried to build a follow-up experiment which failed and was not published because the failure to replicate will usually be assumed to be due to some mistake, and even carefully finding p > threshold is not very publishable.
A large amount of published results that are wrong is definitely something science can live with: we have to trade off Type I against Type II. But we should value accuracy: if we report something as being very, very unlikely if chance was at play, and it turns out that in fact (1) it'd be very likely even if the null hypothesis holds and (2) in fact even if P(D|H0) is low, P(H0|D) might be high... then what's the point in writing up all those fancy statistical analyses anyway? At that point significance testing becomes more of religious ritual and should either be discarded entirely or be amended.
If the p-values were accurate and averaged around 0.05, ~95% of results should be reproducible.
That only 36% were points to deep, fundamental errors.
No. P-values don't work that way and don't mean what you think they mean. Read OP or heck, any of the classics like "Why most published research findings are false" http://dx.plos.org/10.1371/journal.pmed.0020124 http://dx.plos.org/10.1371/journal.pmed.0020124
(36% may or may not be bad, but you can't know without additional stuff like power or prior probability of hypotheses being true; p-values have no intuitive meaning and aren't an answer to any question that people are asking, which is a major reason why Bayesian approaches can be useful. And from a Bayesian perspective, I find 36% totally unsurprising - if anything, substantially better than I had expected given the gross underpowering of most psych studies, the statistical-significance publication filter, and the dubiousness of most hypotheses.)
A proper rebuttal would show what a p-value actually is and how it differs from what I claimed. Now, since a p-value is exactly what I previously claimed, you obviously can't do that. I'm not even sure what you are arguing against me here.
The p-value is the chance of a false positive. But you don't know what the rate of true positives is, or the rate of false negatives.
In a world where there are only false positives and true negatives, and people publish all positive and negative results, then reproduction of a paper should be 95%.
But the reproduction rate when there actually is an effect is not 95%. Depending on sample size, I might get a true positive 20% of the time and a false negative 80% of the time, or I might get a true positive 99.8% of the time and a false negative .2% of the time.
So the average reproduction rate, where an effect actually exists, can be almost any number between 5 and 100. There is no reason to assume it will be 95%.
So the average reproduction rate, where some effects are real and some are imaginary, will almost certainly not be exactly 95%, and that is not a problem in and of itself.
(And when you talk about an average p-value of .05, that sounds like only publishing positive results, which is blatantly going to fail reproduction. 100 false hypotheses -> 5 publications, all false positives -> 5% reproduction rate)
Let me join the club of people claiming that you don't understand p-values.
It's not clear if you are talking about the rate of reproduction for a subset of the possible experimental outcomes (those rejecting the null at the alpha=5% level) or for the whole set.
When the null hypothesis is true (remember that there are fields where this is the norm, v.g. ESP), you would only reproduce (reject again) 5% of the rejections.
Of course you would reproduce (non-reject for the second time) 95% of the non-rejections. The global reproduction rate would be 0.95 x 0.95+0.05 x 0.5=0.905 (90.5% doesn't look ~95% either).
When the null hypothesis is not true, the probability of reproducing (in either sense) the result of a test depends on the effect size.
If the effect is huge, the test will be rejected with probability ~100% and the result will be reproduced with probability ~100%.
Or maybe you mean by reproducing "getting a lower p-value" in the the second trial? If the null is true, the probability of getting p-value2<p-value1 is precisely p-value1. If the null is not true, it will depend on the effect size. If you assume the effect size is the observed one, you expect p-value2 to be smaller than p-value1 with probability 50%.
> I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05.
What do you mean with "measure a p-value"? You make your observation, calculate a statistic (a function of the observation), and look at the distribution of that statistic under the null. The p-value is, by definition, the percentile of the value you got in that distribution (which might or might not be the actual distribution).
You want to check if a die is loaded to yield 6 more often than it should. The null hypothesis is that the die is fair. You can calculate the distribution for the number of 6's in 3 rolls (0: 58%, 1:35%, 2: 7%, 3: 0.5%). You roll the die three times, you get three 6's. The p-value is 0.005. Do you agree? The p-value is 0.005 whether the die is fair (the null hypothesis is true) or loaded. Do you agree?
> Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science.
Scientific experiments are usually about rejecting the null hypothesis. For example, the null hypothesis might be that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001, do you think they calculated it properly?). In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated" and not "repeat the experiment and get a result consistent with the null hypothesis".
According to your description of the limits of p-value analysis, the only conclusion that physicists should get out of the experiment is that if they do it again they should expect to get results consistent with the null hypothesis (i.e. no Higgs boson) with 95% probability. But they see it as evidence that the null hypothesis is false and the Higgs boson real.
Ok, so you're thinking about a random variable which converges to some value when the null hypothesis is true. This is fine, but it has nothing to do whatsoever with p-values.
Let me say that your notation is not very appropriate. It makes no sense to say that P(X|H) converges to 1. If you expect X to converge to C if the null hypothesis is true, you can simply say X->C. A proper notation involving probabilities would be P(|X-C|>epsilon)->0 for any positive epsilon (convergence in probability) or maybe P(X->C)=1 (convergence almost surely).
Taking as you suggest X=(#tails/#heads), you expect that X->1 if the coin is fair (I'm not sure why you find this is not a well-defined null hypothesis, but I don't really care). However, P(X)<1 for every X. In fact, P(X=1)->0 as the number if trials increases (X will get closer to 1 on average, but getting exactly 1 will get more and more unlikely).
As I said, you're free to prefer your converging statistics and your well-defined null hypothesis. But you should be aware that people are talking about something completely different when discussing things like the 1e-7 p-value in the Higgs boson discovery or the reproducibility of statistically significant results.
EDIT: Another example, maybe better-defined: a random variable distributed (under the null hypothesis) x~Normal(mu=0,sigma=1). Let's say you take N samples (I let you choose the number, so I don't pick one which is not good enough).The statistic is the mean X=(x_1+x_2+..+x_N)/N. If the null hypothesis is true, X->mu=0. You get X=1/sqrt(N). What's your "p-value" in that case?
>Ok, so you're thinking about a random variable which converges to some value when the null hypothesis is true. This is fine, but it has nothing to do whatsoever with p-values.
Well, the variable itself doesn't, just the observed value of P(X|H). X can be any random variable, but typically it will need to be transformed to have a normal distribution about 0 with a standard deviation of 1 (since this is what the typical null-hypothesis predicts).
To effectively use p-value analysis, it is typically assumed that your null-hypothesis predicts that your observations will be normally distributed with a mean of 0 and a standard deviation of 1. The total count of heads observed will not be distributed that way. Neither will the probability of a particular sequence (what your example seemed to be calculating). I say your null hypothesis is not well-defined because the term 'fair' remains undefined (though we could guess at the meaning) and in fact makes no predictions about the world. You need to apply transformations to your random variable so that it will appear normally distributed about 0 with a standard deviation of 1 if the hypothesis is true.
>Let me say that your notation is not very appropriate. It makes no sense to say that P(X|H) converges to 1.
My notation is perfectly appropriate. X is a random variable and a random variable is the only thing that can go there (if you are doing p-value analysis). X is not assumed to be uniform or simple (although it certainly could be). P(T>T(X)|H) can be replaced with P(Y|H) every time (Y = T>T(X)).
>As I said, you're free to prefer your converging statistics and your well-defined null hypothesis. But you should be aware that people are talking about something completely different when discussing things like the 1e-7 p-value in the Higgs boson discovery or the reproducibility of statistically significant results.
I'm glad that we finally agree on this (although I dispute that anyone working on the Higgs boson discovery disagrees with me). One of my first claims was that others may not be calculating true p-values, but may calculate something and call it 'p-value' and then think that it means something it does not. In fact, this entire topic even links to an article in a prominent publisher claiming the same.
Do you think it is purely coincidental that the figures I showed you from the Higgs experiment show the lines converging towards only two different numbers: 1 and 0?
Edit: You'll have to give me some time on your edit. It's not something I typically calculate and I have other business to attend to today.
Ok, so maybe your definition does correspond to a p-value after all. It's hard to say as you have refused to discuss concrete cases (like the fair coin or the loaded die, which are standard examples to introduce p-values). But if you're actually calculating a p-value then it won't behave as you expect. It won't converge to anything (edit: if the null hypothesis holds). P-values are by definition uniformly distributed when the null hypothesis is true. If your "p-value" is not, then it's not a p-value. It really is that simple. Or maybe everyone else is using the wrong "p-values" and yours are the real thing. You can believe it if you want.
Please disregard my previous questions, I see no point in continuing this discussion. But you might want to read a bit more about p-values: you won't find anyone (I hope!) sharing your point of view. Once you understand what the p-value is, and what it is not, you might indeed conclude that they are entirely useless. Of course it's your right to avoid learning what p-values really are, and keep the faith. It's your choice.
Do you think it's purely coincidental that this this figure https://atlas.web.cern.ch/Atlas/GROUPS/PHYSICS/CONFNOTES/ATLAS-CONF-2012-170/fig_03.png https://atlas.web.cern.ch/Atlas/GROUPS/PHYSICS/CONFNOTES/ATL... includes the sigma=0 (null hypothesis) line at 0.5 and not at 1? (Hint: the expect value of the p-value under the null hypothesis is 0.5.) (That's a rhetorical question: I already know it's because this is not a well-formed experiment or something.)
>P-values are by definition uniformly distributed when the null hypothesis is true.
Where are you getting this from? When the null hypothesis is true, the p-value should be 1. This follows directly from the definition. If the p-value is not 1 and the null hypothesis is in fact true, your experiment or calculations are wrong. You might also just have the wrong null hypothesis (eg. sensors have more noise than assumed).
I'm willing to accept that in some designs, the first observation of p-value may be uniformly distributed over (0,1], but, as additional observations are made, the value should converge to 0 or 1. What would be the purpose, or usefulness of p-value being uniformly distributed if the null-hypothesis is true? It's much simpler to design things to converge to a single number.
Edit: I have also considered that p-value could be uniformly distributed if the null-hypothesis is false (where you claimed true). I don't know the answer to that.