4 ms·
It's highly relevant to the discussion (interestingly titled: "P-value not as reliable as many scientists assume"), which is entirely about p-value, as I am des
by jsprogrammer 11y ago
It's highly relevant to the discussion (interestingly titled: "P-value not as reliable as many scientists assume"), which is entirely about p-value, as I am describing to you the limits of p-value analysis.
That some perform calculations that are not p-value and call them p-value is not exactly my problem to solve. That others perform meta-analyses with numbers that others call p-values, but which aren't actually p-values isn't really my problem either.
I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05.
Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science.
If you expect 95% of observations (p = 0.05) to be consistent with previous findings, but only 36% are...you did not calculate a valid p-value (or are now testing something other than your hypothesis).
- kgwgk 11y ago> I'll say it again. If you correctly measure (exercise left to reader) a p-value of 0.05, that measurement explicitly means that you expect that 95% of your future observations to be consistent with the hypothesis which you used to determine that p-value of 0.05. What do you mean with "measure a p-value"? You make your observation, calculate a statistic (a function of the observation), and look at the distribution of that statistic under the null. The p-value is, by definition, the percentile of the value you got in that distribution (which might or might not be the actual distribution). You want to check if a die is loaded to yield 6 more often than it should. The null hypothesis is that the die is fair. You can calculate the distribution for the number of 6's in 3 rolls (0: 58%, 1:35%, 2: 7%, 3: 0.5%). You roll the die three times, you get three 6's. The p-value is 0.005. Do you agree? The p-value is 0.005 whether the die is fair (the null hypothesis is true) or loaded. Do you agree? > Making future observations that are consistent with a known hypothesis is exactly what reproducibility refers to within the context of science. Scientific experiments are usually about rejecting the null hypothesis. For example, the null hypothesis might be that there is no Higgs boson and the peak observed in the LHC data is just noise. They made their analysis and rejected the null hypothesis (p-value less than 0.000001, do you think they calculated it properly?). In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated" and not "repeat the experiment and get a result consistent with the null hypothesis". According to your description of the limits of p-value analysis, the only conclusion that physicists should get out of the experiment is that if they do it again they should expect to get results consistent with the null hypothesis (i.e. no Higgs boson) with 95% probability. But they see it as evidence that the null hypothesis is false and the Higgs boson real.
- jsprogrammer 11y agoMeasuring a p-value is equivalent to calculating a p-value (ie. calculate the conditional probability P(X|H)). I don't really agree that your die experiment is well-formed. For one, you are grossly under-sampling. It's known a priori that there are at least six possible outcomes, yet you are only considering three rolls, so you don't even have the possibility of observing each distinct value even once. The p-value of a well-formed experiment should converge towards a fixed value as more observations are made. You will experience variance in the computed value due to the inherently discrete nature of experimentation. This will be especially pronounced for the first observations that are made. I do not know if the Higgs boson experiment is well-formed. If it is well-formed and their null-hypothesis is true, their p-values will trend towards 1. If their null-hypothesis is not true then the p-values do not mean much and will trend towards 0. >In this context, reproducibility means "finding the Higgs boson again if the experiment is repeated" "The Higgs boson exists" is not a valid hypothesis. Usually the null-hypothesis is "the explanation is measurement/background noise". Since that is really the only valid null-hypothesis, it is most likely what they are using.
- kgwgk 11y agoIt's clear that you have your own concept of a p-value, which is quite different from the one used by all the other people (including the proper interpretation and the usual misinterpretations). You disagree with all the provided examples, but you have not given any concrete example of how the p-value would be used in a "well-formed experiment" (another concept that seems unique to you). Of course you're free to redefine concepts as you please, if it makes you happy or it is useful to you in any other way.
- jsprogrammer 11y agoI have redefined nothing. Go pull the the Higgs data; it will be as I say. Go read how to form experiments and calculate p-values, nothing will be substantially different than what I have said here.
- 11y ago