3 ms·
Conversely though, it is often impossible to obtain the original code to replay and identify differences once that step is reached without some sort of strong i
by BadInformatics 6y ago
Conversely though, it is often impossible to obtain the original code to replay and identify differences once that step is reached without some sort of strong incentive or mandate for researchers to publish it. When the only copy is lost in the now-inaccessible home folder of some former grad student's old lab machine, there is a strong disincentive to try replicating at all because one has little to consult on whether/how close the replicated methods are to the original ones.
- dllthomas 6y agoAnd so we find ourselves in the same situation as the rest of the scientific process, throughout history. When I try to replicate your published paper and I fail, it's completely unclear whether it's "your fault" or "my fault" or pure happenstance, and there's a lot of picking apart that needs to be done with usually no access to the original experimental apparatus and sometimes no access to the original experimenters. The fact that we can have that option is an amazing opportunity that a confluence of attributes of software (specificity, replayability, easy of copying) afford us. Where we are not exploiting this like we could be, it is a failure of our institutions! But it is different-in-kind from traditional reproducibility.
- BadInformatics 6y agoOf course, but the flip side is that same confluence of attributes has also exacerbated issues of reproducibility. Just as science and the methods/mediums by which we conduct/disseminate it have changed, so too should the standard of what is considered acceptable to reproduce. This is especially relevant given how much broader the societal and policy implications have become. More concretely, it is 100% fair (and I might argue necessary) to demand more of our institutions and work to improve their failures. I'm sure many researchers have encountered publications of the form "we applied <proprietary model (TM)> (not explained) to <proprietary data> (partially explained) after <two sentence description of preprocessing> and obtained SOTA results!" in a reputable venue. Sure, this might be even less reproducible 200 years ago than now, but the authors would also be less likely to be competing with you for limited funding! Debating about the traditional definition of reproducibility has its place, but we should also be doing as much as possible to give reviewers and replicators a leg up. This is often flies in the face of many incentives the research community faces, but shifting blame to institutions by default (not saying you're doing this, but I've seen many who do) is taking the easy road out and does little to help the imbalanced ratio of discussion:progress.