4 ms·
The two main problems in academia are that a) few researchers have formal training in best practices of software engineering, and that b) time pressure leads to
by slhck 6y ago
The two main problems in academia are that a) few researchers have formal training in best practices of software engineering, and that b) time pressure leads to "whatever worked two minutes before submission deadline" becoming what is kept for posteriority.
When I started working as a full-time researcher, I had come from working two years in a software shop, only to find people at the research lab having never used VCS, object-oriented programming, etc. Everyone just put together a few text files and Python or MATLAB scripts that output some numbers that went into Excel or gnuplot scripts that got copy-pasted into LaTeX documents with suffixes like "v2_final_modified.tex", shared over Dropbox.
Took a long time to establish some coding standards, but even then it took me a while to figure out that that alone didn't help: you need a proper way to lock dependencies, which, at the time, was mostly unknown (think requirements.txt, packrat for R, …).
- justinmeiners 6y agoDon't you think docker, dependencies, unit test frameworks, etc actually increase the need for ongoing maintenance as opposed to spitting out some C files or python scripts which last "forever"?
- slhck 6y agoI don't think so. The source code is the same but there's now metadata that helps in setting up the same environment again, even years later. You still have the original code in case, e.g. Docker is no longer available. For instance, if you just have a Python script importing a statistical library, what version are you going to use? Scipy had a pretty nasty change in one of its statistical functions, changing the outcome of significance tests in our project. Depending on which version you happened to have installed it'd give you a positive or negative result.
- justinmeiners 6y agoIt makes sense that having more information is better than less. I would argue that they should use no dependencies to avoid this problem entirely, or download them and include them as source in the project, or at least include a note of which version of a major library they used in a README or comment. I think this is what is often done in practice currently. Perhaps as you are saying, docker is just a stable way to document this stuff formally. But it is a large moving part that assumes a lot of stuff is still on the internet. What if the docker hub image is removed or dramatically changed? What if that OS package manager no longer exists? It just doesn't seem like our software is getting more longevity, but less. I don't know why we would bring that extra complexity to academic research if the goal is longevity.
- tanilama 6y agoNo. Python/C files didn't work in a vacuum. They need dependencies, that is the point of Docker after all. Capture all necessary dependencies into a single image.
- justinmeiners 6y ago> Python/C files didn't work in a vacuum They do if you use the standard library (which for python is quite extensive), and copy any dependencies into your own source, as if they are your own. By "in a vaccum" we can mean if python o is installed, it will work. > Capture all necessary dependencies Docker doesn't capture any dependencies. They still exist on the internet. It just captures a list of which ones to download when you build the image. Do you think software we write now has more longevity than older software that uses make or a shell script?
- konjin 6y agoLinus is still able to run the first ./a.out he built on pre-0.1 Linux. I can build a docker file from 5 years ago because all the links are dead.
- hobofan 6y agorequirements.txt is not a lockfile