3 ms·
I'll start by saying that I don't have a good answer to this problem. But, I do have some thoughts. I recently reviewed a paper and was extremely happy to se
by ylem 8y ago
I'll start by saying that I don't have a good answer to this problem. But, I do have some thoughts. I recently reviewed a paper and was extremely happy to see that they included a Jupyter notebook and their Keras model along with some testing data. For those of you not in science, this was rather amazing And when the paper is published, it will be part of the supplemental materials, so at least for awhile, readers will be able to play with it (and the authors plan to release their training data once they find a good way to share such a large set of data--again, a limit for academic researchers). So, this is going above and beyond what I've seen in normal practice (I'm in condensed matter physics).
But...there is the problem of context. In their case, they had relatively few dependencies and I was able to get everything to run, but there are more complex environments and ecosystems. Even if I create a docker container, at some point it will no longer run. I think what we can do is try to make it possible for referees and early authors to run our code for a time. We can't hope that 5 or 10 years from now that this will be possible--but hopefully if we document our reduction steps, then if someone really wants to reproduce the work, they can see the flow.
Now, why might this become important? One example is a case of outright fraud. I went to a talk by someone from MD Anderson who gave a talk about a case of fraud at Duke in an oncology study. They saw an amazing result and their colleagues wanted to be able to use the same statistical methodology. They initially tried to work with the original authors, but once they discovered problems with the work, the original author stopped being responsive. They spend an amazing number of man-years trying to reproduce the result and figuring out what went wrong (intentionally and not). This was important because human trials were beginning. If the original source code (and infrastructure) was publicly available, this could have been avoided.
For those that say a mathematical description should be sufficient--I would say, not always. In some cases, the math could be fine, but the implementation could be flawed. Often if you find an error in a previous result, you need to at least make a guess as to what could have gone wrong before. The early days of Monte Carlo simulations sometimes suffered from flaws in the implementations of random number generators even if the over algorithm was fine...
Containerization might solve the problem over the short term (which I would argue is the most relevant time period). But, it won't solve the author's second problem which is maintaining software. Here, I think the problem is a lack of resources--there's not much credit or funding for maintaining scientific software...
- danyx 8y agoFor sharing large training data have a look at https://zenodo.org https://zenodo.org which is run by the CERN people. Up to 50GB is no problem and after that they say just talk to us :).