8 ms·
Landmark study on the replicability of psychological science overturned [pdf] (2016)
- wolverine876 4y agoThe OP is the press release. Here is the paper and links to the original study author's reply, a reply to the reply, etc. https://projects.iq.harvard.edu/psychology-replications https://projects.iq.harvard.edu/psychology-replications
- ethanbond 4y ago> if you read the reports carefully, as we did, you discover that many of the replication studies differed in truly astounding ways—ways that make it hard to understand how they could even be called replications.” > As an example, Gilbert described an original study that involved showing White students at Stanford University a video of four other Stanford students discussing admissions policies at their university. Three of the discussants were White and one was Black. During the discussion, one of the White students made offensive comments about affirmative action, and the researchers found that the observers looked significantly longer at the Black student when they believed he could hear the others’ comments than when he could not. > So how did they do the replication? With students at the University of Amsterdam!” Gilbert said. “They had Dutch students watch a video of Stanford students, speaking in English, about affirmative action policies at a university more than 5000 miles away.” In other words, unlike the participants in the original study, participants in the replication study watched students at a foreign university speaking in a foreign language about an issue of no relevance to them! I’ve long said it’s much more useful to think of the replication crisis as a generalization crisis, and this is a good case in point. The replication (and failure to replicate in Amsterdam) doesn’t tell us the original study was wrong, but tells us that something in the context was actually part of the subject all along. In this case US/Dutch differences in current social hot topics, or in historical slavery, or maybe it only replicates for students who you believe you might run into in class. In any case, per usual, the solution is more science and more thoughtful study, not throwing in the towel and declaring the sciences conquered by “the replication crisis.”
- uniqueuid 4y agoI agree that a nuanced approach is the right way. But many disciplines, in this case psychology, did have the replication crisis coming. Cohen and Gigerenzer have warned about power for literally decades but their warnings were not heeded by a significant part of the field. So a stark message may in fact have been necessary. By the way, Daniel Kahneman basically wrote a "this is a problem, deal with it NOW" letter to some colleagues in 2012 [1]. Very telling and an insightful read. [1] http://www.decisionsciencenews.com/?p=3634 http://www.decisionsciencenews.com/?p=3634
- jvanderbot 4y agoHere's what he suggested: * Assemble a group of five labs, where the leading investigators have an established reputation (tenure should perhaps be a requirement). Substantial labs with several students are the most desirable participants. * Each lab selects a recent demonstration of a priming effect, which they consider robust and most likely to replicate. * The board makes a public commitment to these five specific effects * Set up a daisy chain of labs A-B-C-D-E-A, where each lab will replicate the study selected by its neighbor: B replicates A, C replicates B etc. * Have the replicating lab send someone to see how subjects are run (hence the emphasis on recency – the experiments should be in the active repertoire of the original lab, so that additional subjects can be run with confidence that the same procedure is followed). * Have the replicated lab send someone to vet the procedure of the replicating lab as it starts its work * Run enough subjects to guarantee power (probably more than in the original study) * Use technology (e.g. video) to ensure that every detail of the method is documented and can be copied by others. * Pre-commit to publish the results, letting the chips fall where they may, and make all data available for analysis by others. I hope that psychology is a field where it is possible to fund such a protocol from grants or from internal school funds. However, that protocol seems to me to be laborious, expensive, and not at all "easy to get".
- wolverine876 4y ago> But many disciplines, in this case psychology, did have the replication crisis coming. The point of the OP is that there isn't a replication crisis. Should we be assuming, for purposes of this discussion, that there is one?
- uniqueuid 4y agoFirst, this is [2016]. Second, as ethanbond has commented, a key question in replication is how strictly one should interpret study designs and how much agreement one would expect. I do believe though that there is a true replication crisis, not just because the situational context is different (i.e. generalization) but also because a lot of the original studies were unintentionally or intentionally wrong, had way too little power, i.e. noisy measurements, small N and publication bias. Recent advances in the field of replicability include specification curve analyses, for example, and those have been able to show more conclusive evidence of non-replicability because the published results were unbelievably special (i.e. one out of a tiny fraction of significant results).
- wolverine876 4y ago> a lot of the original studies were unintentionally or intentionally wrong, had way too little power, i.e. noisy measurements, small N and publication bias. That seems much too obvious for all these scientists to overlook, including the OP. Do the authors of the original replication paper say that? > Recent advances in the field of replicability include specification curve analyses, for example, and those have been able to show more conclusive evidence of non-replicability because the published results were unbelievably special (i.e. one out of a tiny fraction of significant results). Would you have links to this research? Thanks.
- uniqueuid 4y agohttps://www.nature.com/articles/s41562-020-0912-z https://www.nature.com/articles/s41562-020-0912-z Simonsohn and Simmons have a couple of great other papers as well. Also: Sure, many of the problems in sciences are obvious, just as many of the problems in other parts of society are. That doesn't mean that it's rational for most people to care about them. Plus there are some hard problems very adjacent to easy problems, and that's an invitation to throw them all in a bucket and not care. Examples: Low N is easy to recognize, but near impossible to change when studying rare diseases. Non-linear phenomena in linear statistical tests are easy to spot, but incompleteness of models (key valuables omitted) or specification errors in structural equation models are hard to spot (need lots of domain expertise). [edit]: If you want to do a spec curve analysis, try this easy package https://cran.r-project.org/web//packages/specr/index.html https://cran.r-project.org/web//packages/specr/index.html
- deleted 4y ago[deleted]
- wolverine876 4y agoI think one essential question of such issues is that why certain populations are so inclined to believe a specific result. And in case you think I'm talking about the field of psychology, I'm (edit: also or especially) talking about populations like HN, which consistently finds fault with almost all scientific resesarch (except research which finds fault with other research). Also, a key overlooked result of the original replication paper was that the results of the original studies were reproduced, but had lower power than the original studies reported (and of course the OP calls that conclusion into question).
- spacemanmatt 4y agoIs that one essential question answered by Dunning and Kruger's research? I keep going back to it (and not in a trivial way) because it really does shed so much light on behaviors informed (or not) by learned experience.
- wolverine876 4y agoWould you kindly expand on that? What research? What answer does it give? Any links? Thanks.
- spacemanmatt 4y agoObviously I can only nutshell it here for you. Their research delved into self-estimates of expertise compared with objective measurement of expertise. To an extent, they proved the old lamentation that fools are confident and wise men are cautious. I recommend starting here: https://en.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effect https://en.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effect
- itsdrewmiller 4y agoThis is too perfect - that effect also is likely not real: https://statmodeling.stat.columbia.edu/2021/10/12/can-the-dunning-kruger-effect-be-explained-as-a-misunderstanding-of-regression-to-the-mean/ https://statmodeling.stat.columbia.edu/2021/10/12/can-the-du...
- kryptiskt 4y agoThe authors of the study replied: http://retractionwatch.com/2016/03/07/lets-not-mischaracterize-replication-studies-authors/ http://retractionwatch.com/2016/03/07/lets-not-mischaracteri... In there are links to other critics of it.
- Isinlor 4y agoAssuming that everything in this pdf is correct, the conclusion should be that the "The Reproducibility Project: Psychology" does not provide evidence one way or another. There may be replication crisis, but the study could not tell the difference between replication crisis and 100% reproducibility. But today there are a lot more studies that find issues with replicability: - Many Labs 2: Investigating Variation in Replicability Across Samples and Settings: https://journals.sagepub.com/doi/10.1177/2515245918810225 https://journals.sagepub.com/doi/10.1177/2515245918810225 - Reproducible brain-wide association studies require thousands of individuals: https://www.nature.com/articles/s41586-022-04492-9 https://www.nature.com/articles/s41586-022-04492-9 - Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015: https://www.nature.com/articles/s41562-018-0399-z https://www.nature.com/articles/s41562-018-0399-z - Raise standards for preclinical cancer research: https://www.nature.com/articles/483531a https://www.nature.com/articles/483531a - Reproducibility Project: Cancer Biology: https://www.cos.io/rpcb https://www.cos.io/rpcb - A Survey on Data Reproducibility in Cancer Research Provides Insights into Our Limited Ability to Translate Findings from the Laboratory to the Clinic: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0063221 https://journals.plos.org/plosone/article?id=10.1371/journal...
- popcube 4y agolet me indicate your second item is a research that use the data and money more hundred times than most of old research... why you surprised to they have new find? their resources are nobody dare to dream in past.
- tinalumfoil 4y agoLet's say Galileo drops two spheres of independent mass off the Leaning Tower of Pisa, and publishes a study saying they both fall with the same acceleration and concluding this is true of all objects dropped on Earth, assuming they're heavy enough to ignore air resistance. Is it evidence against this study if you perform the same experiment from the Empire State Building and get a different result? Yes, because it matters that the study's conclusion wasn't specific to one building. It's a bit of a low blow to say changing the country the study was performed in invalidates it as a reproduction, if the study itself made conclusions broader than a single country.
- andai 4y agoModafinil increases histamine release in the anterior hypothalamus of rats... in Japan!
- notahacker 4y agoOn the other hand, if you perform a slightly different experiment from the Empire State Building and get different results whilst someone else replicates it perfectly in Pisa, is it justified to ignore the Pisa replication altogether, write off gravity as a theory and add it to your meta-analysis of replication crises in physics, or should you in fact be looking at how the differences in experiment might affect the results? (Maybe the Empire State Building is too tall or not Italian enough for gravity to take effect, or maybe the problem is you're dropping things that glide...). Especially when most of the failed "replications" were things the original study authors said up front weren't valid replications of what they were studying. It's possible to predict in advance that people from a different culture will have a different reaction than Stanford students to Stanford students talking about a hot button issue on their campus (it is also not surprising culture has more impact on psychology than physics!). And even pure ad-hoc explanations of the effect of differences in experimental design (OK, we think the problem might not be the Empire State Building, it might be that they were dropping gliders!) after an unexpected replication fail are more helpful testable hypotheses for science than writing off the original theory altogether. (Though sometimes good hypotheses include "the construction of the original experiment made this mistake, which could be tested by replicating with and without this source of sampling bias ...")
- robocat 4y agoThe OSC responded to this criticism here: https://www.science.org/doi/10.1126/science.aad9163 https://www.science.org/doi/10.1126/science.aad9163 Across multiple indicators of reproducibility, the Open Science Collaboration (1) (OSC2015) observed that the original result was replicated in ~40 of 100 studies sampled from three journals. Gilbert et al. (2) conclude that the reproducibility rate is, in fact, as high as could be expected, given the study methodology. We agree with them that both methodological differences between original and replication studies and statistical power affect reproducibility, but their very optimistic assessment is based on statistical misconceptions and selective interpretation of correlational data. Gilbert et al. focused on a variation of one of OSC2015’s five measures of reproducibility: how often the confidence interval (CI) of the original study contains the effect size estimate of the replication study. They misstated that the expected replication rate assuming only sampling error is 95%, which is true only if both studies estimate the same population effect size and the replication has infinite sample size (3, 4). OSC2015 replications did not have infinite sample size. In fact, the expected replication rate was 78.5% using OSC2015’s CI measure (see OSC2015’s supplementary information, pp. 56 and 76; https://osf.io/k9rnd https://osf.io/k9rnd). By this measure, the actual replication rate was only 47.4%, suggesting the influence of factors other than sampling error alone. Within another large replication study, “Many Labs” (5) (ML2014), Gilbert et al. found that 65.5% of ML2014 studies would be within the CIs of other ML2014 studies of the same phenomenon and concluded that this reflects the maximum reproducibility rate for OSC2015. Their analysis using ML2014 is misleading and does not apply to estimating reproducibility with OSC2015’s data for a number of reasons. First, Gilbert et al.’s estimates are based on pairwise comparisons between all of the replications within ML2014. As such, for roughly half of their failures to replicate, “replications” had larger effect sizes than “original studies,” whereas just 5% of OSC2015 replications had replication CIs exceeding the original study effect sizes. Second, Gilbert et al. apply the by-site variability in ML2014 to OSC2015’s findings, thereby arriving at higher estimates of reproducibility. However, ML2014’s primary finding was that by-site variability was highest for the largest (replicable) effects and lowest for the smallest (nonreplicable) effects. If ML2014’s primary finding is generalizable, then Gilbert et al.’s analysis may leverage by-site variability in ML2014’s larger effects to exaggerate the effect of by-site variability on OSC2015’s nonreproduced smaller effects, thus overestimating reproducibility. Third, Gilbert et al. use ML2014’s 85% replication rate (after aggregating across all 6344 participants) to argue that reproducibility is high when extremely high power is used. This interpretation is based on ML2014’s small, ad hoc sample of classic and new findings, as opposed to OSC2015’s effort to examine a more representative sample of studies in high-impact journals. Had Gilbert et al. selected the similar Many Labs 3 study (6) instead of ML2014, they would have arrived at a more pessimistic conclusion: a 30% overall replication success rate with a multisite, very high-powered design. That said, Gilbert et al.’s analysis demonstrates that differences between laboratories and sample populations reduce reproducibility according to the CI measure. Also, some true effects may exist even among nonsignificant replications (our additional analysis finding evidence for these effects is available at https://osf.io/smjge https://osf.io/smjge). True effects can fail to be detected because power calculations for replication studies are based on effect sizes in original studies. As OSC2015 demonstrates, original study effect sizes are likely inflated due to publication bias. Unfortunately, Gilbert et al.’s focus on the CI measure of reproducibility neither addresses nor can account for the facts that the OSC2015 replication effect sizes were about half the size of the original studies on average, and 83% of replications elicited smaller effect sizes than the original studies. The combined results of OSC2015’s five indicators of reproducibility suggest that, even if true, most effects are likely to be smaller than the original results suggest. Gilbert et al. attribute some of the failures to replicate to “low-fidelity protocols” with methodological differences relative to the original, for which they provide six examples. In fact, the original authors recommended or endorsed three of the six methodological differences discussed by Gilbert et al., and a fourth (the racial bias study from America replicated in Italy) was replicated successfully. Gilbert et al. also supposed that nonendorsement of protocols by the original authors was evidence of critical methodological differences. Then they showed that replications that were endorsed by the original authors were more likely to be replicated than those not endorsed (nonendorsed studies included 18 original authors not responding and 11 voicing concerns). In fact, OSC2015 tested whether rated similarity of the replication and original study was correlated with replication success and observed weak relationships across reproducibility indicators (e.g., r = 0.015 with P < 0.05 criterion; supplementary information, p. 67; https://osf.io/k9rnd https://osf.io/k9rnd). Further, there is an alternative explanation for the correlation between endorsement and replication success; authors who were less confident of their study’s robustness may have been less likely to endorse the replications. Consistent with the alternative account, prediction markets administered on OSC2015 studies showed that it is possible to predict replication failure in advance based on a brief description of the original finding (7). Finally, Gilbert et al. ignored correlational evidence in OSC2015 countering their interpretation, such as evidence that surprising or more underpowered research designs (e.g., interaction tests) were less likely to be replicated. In sum, Gilbert et al. made a causal interpretation for OSC2015’s reproducibility with selective interpretation of correlational data. A constructive step forward would be revising the previously nonendorsed protocols to see if they can achieve endorsement and then conducting replications with the updated protocols to see if reproducibility rates improve. More generally, there is no such thing as exact replication (8–10). All replications differ in innumerable ways from original studies. They are conducted in different facilities, in different weather, with different experimenters, with different computers and displays, in different languages, at different points in history, and so on. What counts as a replication involves theoretical assessments of the many differences expected to moderate a phenomenon. OSC2015 defined (direct) replication as “the attempt to recreate the conditions believed sufficient for obtaining a previously observed finding.” When results differ, it offers an opportunity for hypothesis generation and then testing to determine why. When results do not differ, it offers some evidence that the finding is generalizable. OSC2015 provides initial, not definitive, evidence—just like the original studies it replicated.
- anm89 4y agoIn a completely unscientific manner. I just don't believe this. I just don't believe anything coming anything out of academic psychology in general. It's a bunch of career game players who are engaged in status warfare to be the most prestigious academic psychology game player. I'm not saying I've got proof of this or you should agree with me, it's just my anecdotal read.
- dang 4y agoWe changed the url from https://projects.iq.harvard.edu/files/psychology-replications/files/harvard_press_release.pdf https://projects.iq.harvard.edu/files/psychology-replication... to a non-pdf copy of the same text.