3 ms·
95% of the repeated observations that you make (in the same manner as the observations used to calculate a valid p-value of 0.05) will be consistent with the re
by jsprogrammer 11y ago
95% of the repeated observations that you make (in the same manner as the observations used to calculate a valid p-value of 0.05) will be consistent with the relevant null hypothesis.
What other meaning could there be? The result of an experiment is not a p-value, but a series of observations. Those are what need to be compared.
- kgwgk 11y agoI guess the bit "results should be reproducible" made us think that you were talking about reproducing the previous results (i.e. if the null hypothesis was rejected in the first trial, obtaining again a rejection if the trial was repeated). If I understand your point, you're saying: "If the null hypothesis is true then with 95% probability it won't be rejected. And, independently of the result of the first trial, if we do a second trial and the null hypothesis is true then with probability 95% it won't be rejected". Which seems correct, but you might be overlooking the fact that it's not very interesting and unrelated to the discussion.
- jsprogrammer 11y agoIt's highly relevant to the discussion (interestingly titled: "P-value not as reliable as many scientists assume"), which is entirely about p-value, as I am describing to you the limits of p-value analysis. That some perform calculations that are not p-value and call them p-value is not exactly my problem to solve. That others perform meta-analyses with numbers that others call p-values, but which aren't actually p-values isn't really my problem either. I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05. Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science. If you expect 95% of observations (p = 0.05) to be consistent with previous findings, but only 36% are...you did not calculate a valid p-value (or are now testing something other than your hypothesis).
- kgwgk 11y ago> I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05. What do you mean with "measure a p-value"? You make your observation, calculate a statistic (a function of the observation), and look at the distribution of that statistic under the null. The p-value is, by definition, the percentile of the value you got in that distribution (which might or might not be the actual distribution). You want to check if a die is loaded to yield 6 more often than it should. The null hypothesis is that the die is fair. You can calculate the distribution for the number of 6's in 3 rolls (0: 58%, 1:35%, 2: 7%, 3: 0.5%). You roll the die three times, you get three 6's. The p-value is 0.005. Do you agree? The p-value is 0.005 whether the die is fair (the null hypothesis is true) or loaded. Do you agree? > Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science. Scientific experiments are usually about rejecting the null hypothesis. For example, the null hypothesis might be that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001, do you think they calculated it properly?). In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated" and not "repeat the experiment and get a result consistent with the null hypothesis". According to your description of the limits of p-value analysis, the only conclusion that physicists should get out of the experiment is that if they do it again they should expect to get results consistent with the null hypothesis (i.e. no Higgs boson) with 95% probability. But they see it as evidence that the null hypothesis is false and the Higgs boson real.
- jsprogrammer 11y agoMeasuring a p-value is equivalent to calculating a p-value (ie. calculate the conditional probability P(X|H)). I don't really agree that your die experiment is well-formed. For one, you are grossly under-sampling. It's known a priori that there are at least six possible outcomes, yet you are only considering three rolls, so you don't even have the possibility of observing each distinct value even once. The p-value of a well-formed experiment should converge towards a fixed value as more observations are made. You will experience variance in the computed value due to the inherently discrete nature of experimentation. This will be especially pronounced for the first observations that are made. I do not know if the Higgs boson experiment is well-formed. If it is well-formed and their null-hypothesis is true, their p-values will trend towards 1. If their null-hypothesis is not true then the p-values do not mean much and will trend towards 0. >In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated" "The Higgs boson exists" is not a valid hypothesis. Usually the null-hypothesis is "the explanation is measurement/background noise". Since that is really the only valid null-hypothesis, it is most likely what they are using.