3 ms·
> The vast majority of irreproducible papers aren't detectible as irreproducible at time of publication. They look fine, and many actually are fine. They just d
by ramblenode 3y ago
> The vast majority of irreproducible papers aren't detectible as irreproducible at time of publication. They look fine, and many actually are fine. They just don't reproduce. That's an expected outcome in science.
This is not entirely true. A power analysis is how you determine reproducibility, and researchers should be doing it before they begin collecting data. Reviewers can do it post-hoc with assumptions about the expected effect size (which might come from similar studies). False positives produce inflated effect sizes, so if a result is marginally significant but shows a large effect, that is a good heuristic the result will not reproduce.
- timr 3y ago> A power analysis is how you determine reproducibility All a power analysis does is reduce the chance that the result is a false negative. It doesn't reduce the chance of a false positive. > False positives produce inflated effect sizes Not always. Lots of studies publish as "significant" as soon as they get a p-value just under .05. Inflated effect sizes are certainly a sign that something could be wrong, but it's just one indicator. Regardless, even if you have a power analysis at the conventional threshold of 80%, and a p-value of .05, you're still going to get spurious positive results 5% of the time, and spurious negative results 20% of the time, by definition.
- ramblenode 3y ago> All a power analysis does is reduce the chance that the result is a false negative. It doesn't reduce the chance of a false positive. This is true when we are dealing with an uninformative prior, but published research is known to be biased toward positive results and uncorrected multiple comparisons. This situation leads to small sample studies with high random variance being paradoxically correlated with significant results. High random variance appears as a false large effect size in the published result, so if the power is low when calculated with a smaller (adjusted) effect, there is reason to believe that the p-value is inflated. See e.g. Andrew Gelman's work on small sample studies, garden of forking paths or [0]. > Not always. Lots of studies publish as "significant" as soon as they get a p-value just under .05. Inflated effect sizes are certainly a sign that something could be wrong, but it's just one indicator. Exactly! The implication being the above. [0] https://en.wikipedia.org/wiki/Why_Most_Published_Research_Findings_Are_False https://en.wikipedia.org/wiki/Why_Most_Published_Research_Fi...