3 ms·
It's called multiple testing. You're not guaranteed to get a misleading result, but it would be expected. The definition of the 5% significance level used in p-
by asgraham 8y ago
It's called multiple testing. You're not guaranteed to get a misleading result, but it would be expected. The definition of the 5% significance level used in p-value testing is that given 100 junk datasets (distributed similarly), you would only expect 5 to show as strong an effect.
However, I don't really see what this study has to do with multiple testing. As far as I can tell, they had one dataset, and they tested that, and they got a significant result. Nothing fishy there. Sure the data is noisy, but wouldn't you expect that given how coarse the data is, and how many confounding factors there must be?
[1] https://en.wikipedia.org/wiki/Multiple_comparisons_problem https://en.wikipedia.org/wiki/Multiple_comparisons_problem
- DavidSJ 8y agoPublication bias is a sort of multiple testing at the meta level: different groups of scientists try out different hypotheses — but only one each — and some of them just barely pass the statistical tests and thus get published. Not saying that happened here! I haven’t even read the paper. But it is a valid concern when mining data for patterns, especially when those patterns have a questionable theoretical basis.
- asgraham 8y agoOh absolutely! I totally understand you're not saying that's what's happening here, but just to clarify for other readers, I'd say in this case there are a few reasons we don't need quite that level of suspicion: 1. There is a seemingly-sound theoretical basis for the observation (moreover, there doesn't seem to have been any element of fishing. Sometimes researchers think fishing is OK if they just do it once to avoid multiple testing. But then you certainly have multiple testing at the meta level). 2. This study agrees with at least one other completely independent, reputable study [1] (even different methodologies, apparently) that supported the underlying theory. I do think the problem of multiple testing at the meta level is vastly underrated by the scientific community. Just, in this case, I think there are reasons to reject that hypothesis. [1] https://scholar.princeton.edu/sites/default/files/lwantche/files/beninwkn-qje-final.pdf https://scholar.princeton.edu/sites/default/files/lwantche/f...