9 ms·
I feel like every description of the problem of p-values I read is missing the real issue. And that's that people are using the p-value incorrectly. A fairly s
by thinkmoore 11y ago
I feel like every description of the problem of p-values I read is missing the real issue. And that's that people are using the p-value incorrectly.
A fairly standard definition of the p-value (from wikipedia) says: "In statistics, the p-value is a function of the observed sample results (a statistic) that is used for testing a statistical hypothesis. Before the test is performed, a threshold value is chosen, called the significance level of the test, traditionally 5% or 1% and denoted as α."
What this description is missing though is the crucial importance of the fact that the threshold value is chosen before doing the analysis. And moreover, that the entire analysis plan has been chosen before doing the analysis. Because what the p-value is really telling you is the probability that on repeating the experiment (and its accompanying analysis!) you would see a result as or more extreme than what you observed.
If your experiment comprises "try all of the combinations of variables to see what gives me the best answer", the p-value you compute would need to be some very fancy test that took that into account... and you would see your analysis as having much less statistical power.
For a simple example, look at a statistically rigorous method for dealing with multiple hypothesis testing when you plan it in advance: https://en.wikipedia.org/wiki/Bonferroni_correction https://en.wikipedia.org/wiki/Bonferroni_correction.
Of course p-hacking is bad. The problem isn't frequentist statistics or p-values though, its scientists not understanding the statistics that they use. If you want to use a p-value to help make a decision about a hypothesis, you have to commit to your analysis plan in advance.
edit: Furthermore, p-values were designed to deal with experimental data. If you're doing an observational study, perhaps you should use statistical tools designed for that purpose.
To sum up: when you have people who have no idea what they're doing do statistics, they will do it badly.
- Hermel 11y ago> Because what the p-value is really telling you is the probability that on repeating the experiment you would see a result as or more extreme than what you observed. That's incorrect. If you perform the exact same experiment twice, the chances are exactly 50% that the first result is more extreme than the second one (neglecting equal outcomes). This illustrates nicely how hard it is to properly explain the p-value. Things would be more intuitive if people would report confidence intervals. I prefer "x is with a likelihood of 95% between a and b" a thousand times over "we measured x=c and reject the null hypothesis with 95% probability".
- thinkmoore 11y agoNo, my description is correct, though I left out "assuming the null hypothesis is true." The equal probability you are describing is for the p-value itself being more extreme, not the result.
- Hermel 11y agoYes, adding the clause with the null hypothesis corrects your statement.
- sharnett 11y ago"x is with a likelihood of 95% between a and b" -- that is not what a confidence interval means. Confidence intervals are about as confusing as p-values and don't really solve the problem.
- Retra 11y agoStatistics are just plain confusing all around.
- gjm11 11y ago> the chances are exactly 50% that the first result is more extreme than the second one The chance is 50% if you condition only on the information you have before doing either experiment. But once you've done the first experiment, the chance of a more extreme result given what happened the first time may be much more or much less. (Extreme example: Your experiment consists of rolling ten ordinary 6-sided dice. The null hypothesis is that they're fair dice, fairly rolled, in which case you expect a total not too different from 35 pips. All the dice come up 6. It is not now true that if you run the experiment again, you're as likely to get a more extreme result as you are to get a less extreme one!) > confidence intervals [...] "x is with a likelihood of 95% between a and b" But that isn't what a confidence interval means! A 95% confidence interval [a,b] means "If we ran the experiment lots of times, using the same method of computing the interval [a,b] each time, then in 95% of runs (in the long run) the true value would be in the interval [a,b] obtained on that run". (What you described is what Bayesians call a "credible interval". Of course that interval depends on your prior.)
- goodcanadian 11y agoThe problem with p-values summed up in a comic, one of my all-time favourites, in fact: https://xkcd.com/882/ https://xkcd.com/882/
- PepeGomez 11y agoThis one illustrates it even better: https://xkcd.com/1132/ https://xkcd.com/1132/
- pc2g4d 11y agoIf I understand you correctly, you're saying that p-values are useful when used once to evaluate preselected hypotheses with preselected thresholds. This is a common defense of p-values. Here's what doesn't add up for me: this implies that for many configurations, p-values are invalid. But what was it about preselecting the experimental conditions that really makes them any different? What makes initially chosen hypotheses of higher quality than iteratively discovered, "p-hacked" hypotheses? The need for careful selection of a limited set of variable combinations in advance seems symptomatic that the test being employed is not robust. I'm not convinced that even limited application of p-value significance testing is actually valid.
- thinkmoore 11y agoOne way to think about it is that the p-value is trying to account for sampling and measurement error. Imagine you are a scientist and want to find out whether the hypothesis that men are on average taller than women is true. If you could just exactly measure and take the average of the entire male and female populations you wouldn't need a hypothesis test. Since you can't, you can do an experiment where you take a random sample of men and women. Now, you can do the average in the same way, but you need something to help figure out whether to trust the results. That's where the p-value comes in. The reason you need to be careful to select hypotheses in advance is because in order for statistics to help you account for error you need the noise in the data to be uncorrelated with the result you are trying to assess---which it won't be if you chose the result because it was the one that looked best after you account for the noise.
- unabst 11y ago> To sum up: when you have people who have no idea what they're doing do statistics, they will do it badly. Or rather, when you have people who have a good idea what they're doing, they will do a good job of getting the results they were looking for. And maybe that is why statistics are so popular in political debates.
- thinkmoore 11y agoIn the realm of science, I think it is far more likely that researchers don't realize why what they're doing isn't a good idea. "Never attribute to malice that which is adequately attributed to stupidity." Not that I think researchers are stupid. It's just very hard and very complicated.
- wbillingsley 11y agoA cynical view: the job of a p-value is to make social scientists and psychologists look like they have great publication records. I sometimes think this should be set as a compulsory first exam question in any large-cohort p-value-using discipline: Bob has a class of 100 undergraduate students. He tells them to go and run one experiment each as part of their final year class. Assuming all the data in all the experiments is completely random, what is the probability that at least one student will nonetheless find a "statistically significant effect" (p < 0.05) that can then be written up and get published. rough answer: 99.4% p < 0.05 seems to be chosen as that's a barrier that is not too harsh on the individual researcher (at very high significance levels it would be hard for anyone to run a powerful enough study in order to get published and keep their research job), but p-values are essentially defeated as a filter for genuine effects by the sheer number of experiments being run around the world. In some senses, the answer here is simply to stop brandishing science as a social cudgel ("X is science; how dare you not believe it" or more often "how dare you employ/fund someone who does not believe it, they must be sacked") which puts it on a pedestal of being canonical truth all the time. It is only the fact that science has recently been used more and more as a political stick to beat opponents with (and no, not just in climate and evolution, but right down to things like the best way to teach reading, or whether schools should be regulated or independent) that has meant that people think "science is broken" whenever it turns out to have got a wrong result. No, it's just expected to get it wrong fairly often. And in practice the most common "self-correction" process in science isn't a repeated experiment and published retraction, but just academics reading the paper, thinking "this one's garbage", and not basing their research from it. Publication is not strictly a test of truth -- just a test of methodology and analysis. The test of truth always occurs in the mind of the reader.
- animefan 11y agoI don't think that's a fair characterization of the issue. p-hacking is done almost exclusively by people who completely understand the definition and interpretation of p-values. However these people are also under a lot of pressure to produce positive results (and sometimes, results in a specific direction) and this biases their thinking. The problem is not that people are not aware of the issue of testing multiple hypotheses. The problem is that (1) it's hard to say exactly what your hypothesis is before you've even looked at the data, and (2) it's hard to determine if people are choosing parameters for p-hacking or simply making choices based on their best judgement. >Furthermore, p-values were designed to deal with experimental data. If you're doing an observational study, perhaps you should use statistical tools designed for that purpose. This is simply wrong. p-values are equally relevant in both cases. E.g. I can use p-values to reject the hypothesis that consuming saturated fats is uncorrelated with weight gain, amongst the general population. It sounds like you are reaching beyond your actual expertise in statistics.