7 ms·
When a small sample size produces a big effect, it does not mean you can be more confident in the effect! All things equal, you'd expect confidence intervals /
by closed 9y ago
When a small sample size produces a big effect, it does not mean you can be more confident in the effect!
All things equal, you'd expect confidence intervals / credible regions to be very large.
Another way to say this, is that since the measurement error is large, you expect spurious findings that result from something like p-hacking to be large (they have to be to be significant).
- roganp 9y agoThis. This is a relevant podcast featuring John Ioannidis (author of "Why Most Published Research Findings Are False.") http://www.econtalk.org/archives/2018/01/john_ioannidis.html http://www.econtalk.org/archives/2018/01/john_ioannidis.html
- vanderZwan 9y agoWhat annoys me as a someone not trained in statistics is that these discussions are always abstract in-principle arguments about what proper statistics should be like. Meanwhile, we have a real paper with real p-values, why not discuss the validity of that? > SRT Consistent Long-Term Retrieval improved with curcumin (ES = 0.63, p = 0.002) but not with placebo (ES = 0.06, p = 0.8; between-group: ES = 0.68, p = 0.05). Curcumin also improved SRT Total (ES = 0.53, p = 0.002), visual memory (BVMT-R Recall: ES = 0.50, p = 0.01; BVMT-R Delay: ES = 0.51, p = 0.006), and attention (ES = 0.96, p < 0.0001) compared with placebo (ES = 0.28, p = 0.1; between-group: ES = 0.67, p = 0.04). FDDNP binding decreased significantly in the amygdala with curcumin (ES = −0.41, p = 0.04) compared with placebo (ES = 0.08, p = 0.6; between-group: ES = 0.48, p = 0.07). In the hypothalamus, FDDNP binding did not change with curcumin (ES = −0.30, p = 0.2), but increased with placebo (ES = 0.26, p = 0.05; between-group: ES = 0.55, p = 0.02). I can see that there are a lot of p-values around 0.05, but that the supposed improvements have much lower p-values: (ES = 0.63, p = 0.002) SRT Consistent Long-Term Retrieval (ES = 0.53, p = 0.002) SRT Total (ES = 0.50, p = 0.01) BVMT-R Recall (ES = 0.51, p = 0.006) BVMT-R Delay (ES = 0.96, p < 0.0001) attention What does this imply? Is it a good sign, or does it actually make it more fishy? pythonslange's comment[0] seems to suggest the latter, claiming that none of the other pre-registered tests are mentioned, without any explanation why. (eyeing that "attention" one, if this does hold up under scrutiny, could I expect tumeric-based ADD medication somewhere in the future, without all the nasty side-effects of my current amphetamine-based options?) [0] https://news.ycombinator.com/item?id=16229693 https://news.ycombinator.com/item?id=16229693
- roganp 9y agoHere is the problem: If there is an actual large effect size, then you can detect this in small samples. However, the opposite is not true. In a small sample, spurious large effects are MORE likely, not less likely (outliers have a greater impact in a small sample). See: http://andrewgelman.com/2017/08/16/also-holding-back-progress-make-mistakes-label-correct-arguments-nonsensical/ http://andrewgelman.com/2017/08/16/also-holding-back-progres...
- vanderZwan 9y agoThank you, that is an explanation that I can actually apply to this result.
- closed 9y agoI agree with the other commenter, and Gelman's blog is a good place to start. There are very clear, concrete statistical arguments, but they are difficult to summarize in the comments section of HN. The gist is like this. Imagine that someone had people play 20 different slot machines. Then they went into a private room and looked at the results for each machine. After, they come out with the results of 5 of the machines and say, "look, our slot machines pay out at a higher than chance level!". Do you believe them? I hope not. If only 4 machines had done well, maybe they would have shown only 4, or 3, etc.. They've effectively stacked the deck. On the other hand, suppose someone said they were running an honest experiment with 20 slot machines. How many machines would you expect them to report on?
- vanderZwan 9y agoI can follow your analogy, but it doesn't quite add up. The researchers took a group of people and submitted each of them to an "extensive neuropsychologist test battery". That's not just testing 20 different slot machines, that's testing 20 different designs of slot machines. (Also, each person does each test on their own, which is the equivalent of having them not only play twenty different models, but a unique production per person unit per model) In that light, the claim that all slot machines have a high pay-out chance is obviously suspicious, but would coming back and saying "these five designs have a higher than chance level of winning!" be an incorrect conclusion? If each slot machine types is unique, no. But if one slot machine design is known to have a flaw, and if it shares this flaw with another design, and if that other design does not show the same increased performance, then things get really suspicious. So the question becomes: do we know how strongly correlated the results of these tests typically are? If that is a lot (which I would expect to be true with at least some of these tests), the absence of the other tests is suspicious. If it is low, it might be less of an issue.
- deleted 9y ago[deleted]