8 ms·
This will be quite bad for reproducible science. Publishing bioinformatics tools as containers was becoming quite popular. Many of these tools have a tiny nich
by brutos 6y ago
This will be quite bad for reproducible science. Publishing bioinformatics tools as containers was becoming quite popular. Many of these tools have a tiny niche audience and when a scientist wants to try to reproduce some results from a paper published years ago with a specific version of a tool they might be out of luck.
- hvs 6y agoMaybe they should switch to Github. https://github.com/features/packages https://github.com/features/packages
- toomuchtodo 6y agoOr store the containers in the Internet Archive alongside the paper. They’re just tarballs. Lots of options as long as you're comfortable with object storage.
- captn3m0 6y agoquay is another alternative.
- brutos 6y agoThis still means that tools published in the last few years until now might just be gone soon. The people who uploaded the images might have graduated or moved on and none will be there to save the work.
- icebraining 6y agoSounds like a job for the Archive Team, as long as there's some way to identify the images worth saving.
- maxfan8 6y agoYep, just mentioned it to the Archive Team IRC. We're probably going to selectively archive particular Docker images, although that's a lot of manual labor. If you have any ideas wrt to selecting important images, that'd be great.
- thebouv 6y agoRough idea: maintain an Awesome List of images worth saving, take submissions from public, use that list to automate what to pull?
- maxfan8 6y agoYeah, good idea — I’m not in these fields so it’s difficult for me to judge. Also, it sounds like we should be prioritizing niche images that only a handful of papers use rather than images that people rely upon regularly.
- cosmie 6y agoCouldn't you bootstrap a list by searching/parsing the Archive dataset itself? Searching for A) "docker pull" commands and parsing the text that comes after it based on the command's syntax[1] to extract instructional references to images such as "docker pull ubuntu:latest, and B) Searching for links/text beginning with "https://hub.docker.com/_/" https://hub.docker.com/_/" to identify informational references to image base pages such as (https://hub.docker.com/_/ubuntu https://hub.docker.com/_/ubuntu) [1] https://docs.docker.com/engine/reference/commandline/pull/ https://docs.docker.com/engine/reference/commandline/pull/
- maxfan8 6y agoGood idea! The base images are probably not in danger of being deleted though. The other issue is that (to my knowledge) the amount of papers on IA isn't terribly impressive. I think maybe indexing and going through SciHub will be better since some of these fields slap paywalls in front of their papers. However, that's a pretty large task as well. The other thing is that papers rarely say "to reproduce my work do . . .". Usually the best we've got is a link to a GitHub repo (if that). I'm not sure how effective that strategy will be since it's guaranteed to be an under-count of the docker images we'd need to archive. Perhaps in conjunction with archiving all images that fall under particular search queries, we'd get the best of both worlds. I've you've got ideas, feel free to hop onto efnet (#archiveteam and #archiveteam-bs) (also on hackint) to share your thoughts.
- vegannet 6y agoGitHub storage for docker images is very expensive relative to free: I don’t think it’s a viable solution in this case.
- lstamour 6y agoPublishing containers to GitHub might be free but you have to login to GitHub to download the containers from free accounts, significantly hampering end-user usability compared to Docker Hub, particularly if 2FA authentication is enabled on a GitHub account. As mentioned elsewhere Quay.io might be another alternative.
- qppo 6y agoYou don't need to register an SSH key to download a public repo I thought
- timdorr 6y agoNot an SSH key, but you do need an access token: > You need an access token to publish, install, and delete packages in GitHub Packages. https://docs.github.com/en/packages/using-github-packages-with-your-projects-ecosystem/configuring-docker-for-use-with-github-packages#authenticating-to-github-packages https://docs.github.com/en/packages/using-github-packages-wi...
- qppo 6y ago...but not to download. You can clone a repo and download release artifacts without a PAT. That's only necessary for interacting with the API for actions that need authentication, which would be anything involving mutating a repository.
- akerl_ 6y agoUsing the GitHub Docker Registry requires auth, even just for downloads. https://docs.github.com/en/packages/using-github-packages-with-your-projects-ecosystem/configuring-docker-for-use-with-github-packages https://docs.github.com/en/packages/using-github-packages-wi... GitHub Packages is different from GitHub Releases (and their artifacts) or cloning repos.
- mynameisvlad 6y ago> You need an access token to publish, install, and delete packages in GitHub Packages. Yes, you do.
- chrisandchris 6y agoIt seems you simply have to pull it every 5.99 months to not get it removed. So add all your images into a bash script and pull them every couple weeks using crontab and you‘re fine. On the other side, I see the need for making money and storage/services cannot be free (someone pays somewhere for it - always), but 6 months is not that much for specific usages.
- crazysim 6y agoI'm sure you've cited research older than 5.99 months right? I wish they would grandfather images before this new ToS to not get wiped so that future images would be uploaded to more stable and accepting platforms so images on Docker Hub from research pre-ToS update don't get wiped.
- riffic 6y agowell it sounds like someone's gotta pony up the bucks for a their own image repo, rather than freeload off someone else's storage infra.
- edoceo 6y agoScience Docker Repo as a Service backed by Amazon Glacier and index, one time fee to access?
- salawat 6y agoFull circle achieved. START Run your own stuff on stuff you own. Run your stuff on other people's stuff you rent. This is too expensive to maintain at your rent. Pay us more. Back to run your own stuff on stuff you own. ...and so on, and so forth. And this, ladies and gentlemen, is why anything worth doing is worth actually doing yourself. Nothing is worse than building something conditionally feasible on someone else only to have the rug pulled out from under you by sudden business pivots. But that's the nature of the beast I suppose. I've certainly not found a great way to do it any other way.
- chrisandchris 6y ago
- quotemstr 6y agoWhy? It'll force a shift to a more elegant and general model of specifying software environments. We shouldn't be relying on specific images but specific reproducible instructions for building images. Relying on a specific Docker image for reproducible science is like relying on hunk of platinum and iridium to know how big a kilogram is: artifact-based science is totally obsolete.
- dguest 6y agoI couldn't agree more. The defense of images over instructions to build them has often been "scientists don't work this way", but to me that's either overly cynical or an indication that something is rotting in academic incentive structures.
- CameronNemo 6y ago> rotting I would not say rotting. From my perspective, the academic community has always lagged behind engineering best practices (except in their specific fields).
- eat_veggies 6y agoYou could say the same about distributing docker images for deploying code for non-scientific software as well (and honestly, it may very well be true). But that doesn't change the fact that it's just way easier to skim a paper and pull a docker image than follow every paper's custom build instructions and software stack.
- quotemstr 6y agoWhy would build instructions have to be custom? Making a reproducible image should be as easy as getting a docker image
- dijksterhuis 6y agoThese reproducible instructions you speak of are already present in Dockerfiles. It seems like you're arguing against using docker images, when docker builds solve the very issue you speak of. Correct me if I'm wrong...?
- dguest 6y agoMy field is doing something similar. Reproducible science is definitely a good goal, but reproducible doesn't mean maintainable. Really scientists should be getting in the habit of versioning their code and datasets. Of course a docker container is better than nothing, but I would much rather have a tagged repository and a pointer to an operating system where it compiles. It's true that many scientists tend to build their results on an ill-defined dumpster fire of a software stack, but the fact that docker lets us preserve these workflows doesn't solve the underlying problem.
- MengerSponge 6y agoFYI, and for anyone else still learning how to version and cite code: Zenodo + GitHub is the most feature rich and user-friendly combination I've found. https://guides.github.com/activities/citable-code/ https://guides.github.com/activities/citable-code/
- dguest 6y agoZenodo is great! In theory you could also upload a docker image to Zenodo and give it a DOI, but it doesn't seem to have an especially elegant way to pull this image after the fact.
- wadkar 6y agoThank you for mentioning Zenodo. I really liked how EU funding agencies push for reproducibility/citability of data and code when you submit proposals to them. I haven’t filed any NSF stuff (yet) but didn’t come across any such hard requirements where you had to commit to something like zenodo or else to archive the result of your research work for archiving/citations purposes.
- MengerSponge 6y agoI <3 Zenodo. My societies don't require open data, but that's a generational shift. Also, if you do bio-type research, you can use Data Dryad too!
- dijksterhuis 6y agoSimplest answer is to release the code with a Dockerfile. Anyone can then inspect build steps, build the resulting image and run the experiments for themselves. Two major issues I can see are old dependencies (pin your versions!) and out of support/no longer available binaries etc. In which case, welcome to the world of long term support. It's a PITA.
- OnlyOneCannolo 6y agoYou can also save the image to a file: https://docs.docker.com/engine/reference/commandline/image_save/ https://docs.docker.com/engine/reference/commandline/image_s...
- atomi 6y agoI would recommend running a registry mirror as it's fairly straightforward. https://docs.docker.com/registry/recipes/mirror/ https://docs.docker.com/registry/recipes/mirror/
- OnlyOneCannolo 6y agoThat's still more effort than pushing a tar file to a free public gh repo.
- atomi 6y agoThere is a bit of upfront work, but backups are thereafter automated.
- OnlyOneCannolo 6y agoTrue. I was thinking more for archiving science. Most people in that category would probably rather push to gh or upload to Dropbox than set up a docker registry.
- Nullabillity 6y ago
- Legogris 6y agoAs long as the Dockerfile is released alongside, this should not be an issue. I don't see any valid reason why anyone would upload and share a public docker image but not its Dockerfile and therefore do not pull anything from Dockerhub that doesn't also have the Dockerfile on the Dockerhub page.
- Bedon292 6y agoWhat about when the image that it is based on goes out of date and is pruned too?
- Legogris 6y agoThis is part of why I tend to only use images that only build from a small set of well-established base images like scratch, alpine, debian and occasionally ubuntu. Those base images can also be handled in the same way. For any exception, you can always do the same. A bonus to this is that you no longer have the risks of systems breaking because of Dockerhub or quay.io (which I haven't seen mentioned here yet, btw) being offline.
- jacques_chester 6y agoDockerfiles are not guaranteed to be reproducible. They can run arbitrary logic which can have arbitrary side-effects. A classic is `wget https://example.com/some-dependency/download/latest.tgz` https://example.com/some-dependency/download/latest.tgz`.
- diffeomorphism 6y agoCouldn't journals host the images? Or some university affiliated service, let us call it "dockXiv"? Having the images on dockerhub is more convenient, but as long as the paper says where to find the image this does not seem that bad.
- casept 6y agoThey should be using Nix or similar then. The typical Dockerfile is not reproducible.