3 ms·
A quick recap of the basics (multiple testing appears at the X): Precisely, what a p-value means is: "Assuming the null hypothesis is true, the probability tha
by asgraham 9y ago
A quick recap of the basics (multiple testing appears at the X):
Precisely, what a p-value means is: "Assuming the null hypothesis is true, the probability that I would get this test statistic from the null distribution is p," where p is your p-value.
When you call a p-value "significant" that means you're rejecting some null hypothesis. That null hypothesis entails some null distribution, and what you're saying is: "It is unlikely the test statistic I've calculated comes from that null distribution." Somehow 0.05 has become the p-value of significance, which means we say a test rejects the null hypothesis if the test statistic we calculated had less than a 1 in 20 chance of being drawn from the null distribution.
The t-test is the test the linked paper used. Essentially they compare list of telomere lengths sorted by different habits to see if they're drawn from the same (normal) distribution. For example, the significant comparison was the list of telomere lengths of all subjects who didn't eat any red meat vs. the list of telomere lengths of all subjects who ate "1-2 [red meat] daily". Based on two lists of numbers, you calculate the t-statistic. Under the null hypothesis that the two lists were drawn from the same normal distribution, the probability distribution of the t-statistics is known. If you calculated a t-statistic that falls on the tail of that distribution (meaning, a t-statistic large enough to be very unlikely) then you say: "Aha! The null hypothesis is probably wrong. The two lists come from different distributions. The red meat eaters have significantly longer telomeres."
X Now let's use a really specific example to show why that doesn't scale. Suppose you generate two lists of normally-distributed numbers with the same mean and variance What you're really testing is whether they're drawn from the same distribution which, in this case, we know they're not (we drew them from the same distribution ourselves). Now do that 100 times. If you want (and if you've already downloaded scipy) you can follow along and verify what I'm saying:
from numpy.random import randn
from scipy.stats import ttest_ind
N_tests = 100; N_samples = 500; sig_level = 0.05
results = [ttest_ind(randn(N_samples),randn(N_samples)) for i in range(N_tests)]
is_significant = [result[1] < sig_level for result in results]
print("The proportion of results that are significant is: ", sum(is_significant) / N_tests)
If you run this a few times, you'll see you get ~0.05 every time. If you increase the number of tests (the number of lists you compare), you'll be even more likely to get close to 0.05.
Now, the natural response to this is: "But that's absurd! You've just compared a bunch of random lists! That has nothing to do with a scientific inquiry looking for meaning and relation within a single dataset!" And that's almost true. And there are absolutely well-designed experiments where there are 100 perfectly reasonable hypotheses to test. What this tells you, is that in that perfectly well-designed experiment you will get about 5 significant results (p < 0.05). Hell, you'll probably even get one "highly significant" result (p < 0.01). By definition.
That means in that case it's meaningless to publish your results at the 0.05 significance level. You correct for that with something like Bonferroni correction (which is probably far too conservative) or Tukey's range test [1], both of which say roughly: Given I tried this many tests, how likely was I to get this test statistic under the null hypothesis?
One closing observation: This holds across experiments and papers. What I mean is, if every year you run 20 t-tests, at the 0.05 significance level you can expect to get 1 spurious significant result every year. I suspect most researchers run more than 20 t-tests per year. This is why a single "significant" result mined from a dataset should never be the basis of an academic paper. That's why experimental design and replicability are so important.
[1] I may have failed to respond to your question for the most part, but here's a link! https://en.wikipedia.org/wiki/Tukey%27s_range_test https://en.wikipedia.org/wiki/Tukey%27s_range_test