3 ms·
Unfortunately no. Very much no, even though it's widely believed that that is a good definition/intuition (and used in many places). It's the odds of having th
by fela 10y ago
Unfortunately no. Very much no, even though it's widely believed that that is a good definition/intuition (and used in many places).
It's the odds of having that results due to chance, if the null hypothesis is true[0]. That latter part might sound pedantic, but the whole point is that we don't know how likely the null hypothesis is. If I test wheather the sun has just died[1] and get a p-value of 0.01 it's still very likely that this result is due to change (surely more than 1%)! We need a prior probability (i.e. bayesian statistics) to calculate the probability that the result was due to chance, that is why that partial definition is incomplete and actually very misleading. This point is subtle, but very important to really understand p-values.
Another way to look at it is: if we knew the probability that the result was due to chance we could also just take 1-p and have to probability of there actually being some effect, a probability that hypothesis testing cannot give us.
There is one nice property that hypothesis testing does have (and why presumably it's so widely used): if the idea you are testing is wrong (which actually means "null hypothesis true") you will most likely (1-p) not find any positive results. This is good, this means that if the sun in fact did not die, and use 0.01 as your threshold, 99% of the experiments will conclude that there is no reason to believe the sun has died. So hypothesis testing does limit the number of false positive findings. The xkcd comic is a bit misleading it this regard, yes it does highlight the limitations of frequentist hypothesis testing, but the scenario depicted is a very unlikely one, in 99% of the cases there would have been a boring and reasonable "No, the sun hasn't died".
For an incredibly interesting article about the difficulty of concluding anything definitive from scientific results I highly recommend "The Control Group is out of Control" at slatestarcodex[2].
[0] To be even more pedantic you would have to add "equal or more extreme", and "under a given model", but "if the null hypothesis is true" is by far the most important piece often missing.
[1] https://xkcd.com/1132/ https://xkcd.com/1132/
[2] http://slatestarcodex.com/2014/04/28/the-control-group-is-out-of-control/ http://slatestarcodex.com/2014/04/28/the-control-group-is-ou...
- lisper 10y ago> It's the odds of having that results due to chance, if the null hypothesis is true[0]. Yes, that's right. I don't know why you think this is at odds with what I said. In fact, I clarified this myself a few hours ago in a sibling comment: https://news.ycombinator.com/item?id=13026907 https://news.ycombinator.com/item?id=13026907
- fela 10y ago"what are the odds that the results you observed could have arisen by chance?" If you say it like this it will very easily be misinterpreted. Once your results are in there are two cases: (1) either the null hypothesis is true and you got those results due to chance, or (2) the null hypothesis is false and there was some actual effect outside of the null hypothesis that helped you get the results. Due to this it is very easy to interpret you statement as referring to the probability of (1). Two two following definitions of p-values sound similar but are not: [Correct] The probability of getting the results by chance if the null hypothesis is true P(Results|H0) [Wrong] The probability that you got the results by chance and thus the null hypothesis was actually true P(H0|Results) I'm not saying you didn't get it, but somebody reading what you wrote can very easily be fooled. And there are a lot of dead wrong definitions on the web[0][1][2][3]. [0] https://www.americannursetoday.com/the-p-value-what-it-really-means/ https://www.americannursetoday.com/the-p-value-what-it-reall... [1] https://practice.sph.umich.edu/micphp/epicentral/p_value.php https://practice.sph.umich.edu/micphp/epicentral/p_value.php [2] http://natajournals.org/doi/full/10.4085/1062-6050-51.1.04 http://natajournals.org/doi/full/10.4085/1062-6050-51.1.04 [3] http://www.cdc.gov/des/consumers/research/understanding_scientific.html http://www.cdc.gov/des/consumers/research/understanding_scie...
- RA_Fisher 10y agoThe OP is referring to Fisher's p-value rather than the more common Neyman-Pierson method that you refer to. Fisher's method doesn't have the concept of the null hypothesis. The difference is fascinating and I do believe the Fisherian method is superior if you can't easily replicate.
- fela 10y agoThe definition of p-value is the same independent of method, as far as I can tell the only real difference is that by Neyman–Pearson you just look at whether the p-value is below a threshold, and Fisher looks at p-value as "strength of evidence" valuable in itself. It's still not the probability that your result was due to chance, it's the probability that under the null hypothesis (and you will definitely need one) you would get that value (or more extreme) by chance.