3 ms·
I personally find NHST suspicious, even if the p-value is less than 0.005. It means that, ASSUMING that the null hypothesis is correct, the probability of obser
by justwantaccount 9y ago
I personally find NHST suspicious, even if the p-value is less than 0.005. It means that, ASSUMING that the null hypothesis is correct, the probability of observing the data is less than 0.5%. That's still not zero, though. For example, if I try to decide that native Hawaiians are US citizens, and the null hypothesis is that they are, but since only ~0.2% of total US population is native Hawaiian, NHST would conclude that native Hawaiians aren't US citizens. Conversely, the p-value being high isn't really meaningful either. For example, if I try to decide if a white person is from South Africa, and the null hypothesis is that they are, and since ~9% of the population in South Africa is white, I would fail to reject the null hypothesis, even if the person's actually from Europe or literally any other country in the world. p-values are dependent on the sample size, too, and become more and more sensitive to smaller differences in the ASSUMED population distribution(s), when the effect size can in fact be small. NSHT seems highly dependent on quality null hypothesis that's correct, which you can't really establish since that's what you're trying to find out. So the approach doesn't really conclude anything either way, and just really seems weak as an approach to scientific testing overall.
- majormajor 9y agoI don't really follow. Your examples seem like weird or poorly thought out experiments (well, surveys, not experiments), not anything to do with significance testing. If I wanted to see if native Hawaiians are citizens and didn't have access to law, real data, anything like that, I'd sample from the population of Hawaiians, not citizens as a whole, and then regardless of the null hypothesis being "are citizens" or "aren't citizens" I think the results would be overwhelmingly one-sided enough to work out. And the South Africa example is asking about 1 single person? That not a question of science, then, just a question?
- justwantaccount 9y agoMaybe my example was contrived, but my main argument was that p-values represent P(D|H), or probability seeing your data conditioned on assuming that the hypothesis is correct, and that approach seems inherently flawed to me. If you're doing an experiment, p-values don't say anything about if the hypothesis is true or false, in fact the approach can't theoretically give that information. If you're doing a basic t-test, it goes something like this: 1. You have a null hypothesis you assume to be true 2. Based on this null hypothesis and the central limit theorem, if you theoretically conduct many experiments you expect some summary statistic to have a specific distribution around the 'true' population summary statistic and 'true' population variance, which you don't have so you assume it to be the null hypothesis 3. You compare your experiment's summary statistic to the hypothesized distribution. If the probability of seeing your summary statistic is below some threshold, you say it's unlikely to see this summary statistic again. But the probability of seeing that summary statistic again actually depends on P(H), which NHST doesn't provide any information on. My examples were meant to highlight the nature of conditional probabilities, rather than how real life experiments are conducted.
- Keysh 9y agoYour examples weren't so much contrived as incoherent, to the point where I begin to suspect you don't really understand NHST. "if I try to decide that native Hawaiians are US citizens, and the null hypothesis is that they are, but since only ~0.2% of total US population is native Hawaiian, NHST would conclude that native Hawaiians aren't US citizens." So what is the observation in this case, and what would the corresponding prediction from the null hypothesis be? The observation that "0.2% of the US population is native Hawaiian" has no relation to your claimed null hypothesis at all. The rest of your objection seems like one of those confused arguments trying to rule out basic reductio ad absurdam ("but if X really isn't true, then your arguments about seeing or not seeing the consequences of X have no basis!"). (And the central limit theorem has nothing to do with null hypothesis testing: you can do NHST with completely non-Gaussian statistics.)
- justwantaccount 9y agoYes, you can conduct NHST without the central limit theorem. However, it's used very widely in NHST. Was there anything wrong with what I said about what a typical t-test usually looked like? My lab would use that approach to do molecular biology. You don't seem to understand my argument, so let me rephrase: The example about Native Hawaiians was meant to highlight the nature of conditional probabilities, and p-values are conditional probabilities. Just because p-values are below some threshold doesn't necessarily mean that the null hypothesis is incorrect and therefore should be rejected. Just because p-values values are high doesn't mean that the null hypothesis should fail to be rejected. P-values do not theoretically give that information. It doesn't even represent the probability of observing that value, since it's a conditional probability - as in, the probability of observing that value given that the null hypothesis is true, not the probability of observing that value. If the p-value is below 0.005, can you scientifically, theoretically conclude that the null hypothesis should be rejected? The probability of seeing a Native Hawaiian person given that the person is a US citizen is below the threshold of 0.005, but does that mean the conditioned part (the US citizen thing) should be rejected? Granted, it's hard to relate that example to actual experiments, but my argument is that p-values don't theoretically give any conclusions either way, and trying to make it "scientific" to draw conclusions by introducing thresholds to a conditional probability, no matter how strict, seems inherently flawed. Using it as a single metric among many, to use it as a tool for exploration makes sense to me. Even to make strong suggestions, sure, especially with all the controls RCTs put in. But to make hard conclusions, as in NHST? The approach itself doesn't have the theoretical power to do so.