3 ms·
I try to do something like this with my publications, and encourage others to. My goal is to have the pipeline from raw data to complete figures and manuscript
by cge 16d ago
I try to do something like this with my publications, and encourage others to. My goal is to have the pipeline from raw data to complete figures and manuscript in a repository, with cached data for computationally expensive analysis and for stochastic simulation results, and the option for the user to just use those or run the full pipeline, with or without the same random seeds. I just make clear that the code was run-once code and is going to be messy compared to code refined over time and diverse uses. I generally use Zenodo to a GitHub repo, however, in case GitHub decides to do something bad in the future. Making sure things run far in the future can also be a challenge. Sure, you can use a container: will the base of that container be available in 30 years?
And with that said, for experimental work, this approach does not make things fully reproducible; it only makes the analysis reproducible. There are always factors that influence experiments: research is by definition at the edge of our understanding, and reality has countless variables, including ones no one has thought of, known about or thought important.
- amarcheschi 15d agoYou're goat, I'm trying to reproduce code from a paper and by following their instructions I can't even get packages to install because they conflict
- ErikBjare 14d agoThey probably didn't conflict at the time, see if you tell the resolver to use versions no newer than the publication date (like uv's `exclude-newer` option). I've gotten deps for older research code to install with that one simple trick.
- flopsamjetsam 15d agoCaveat: I am not a researcher (yet), I am moving from coding into science via a new degree, and along the way I am helping troubleshoot bioinformatics pipelines for scientists. I see a lot of reliance on containers to make code always available, and I have the same misgivings as you do. It'll work a few years into the future, but what happens once packages aren't compatible with each other/the base container is upgraded/etc. I've already seen this with older bioinformatics code, which is on old repositories that aren't running anymore, or are very unreliable (but weren't at the time that the code was written). And I'm talking about code that's "only" 15 years old; people will be going to these papers for implementation details long after that point. Zenodo seems like a good step in the right direction. It should be available as long as CERN is going, shouldn't it? And by that stage it should be "too big to fail".