3 ms·
Jupyter notebooks are risky business for reproducible work. No dependency data means they're highly prone to bit rot. Stored results and out-of-order execution
by codebje 7y ago
Jupyter notebooks are risky business for reproducible work. No dependency data means they're highly prone to bit rot. Stored results and out-of-order execution means they're prone to subtle errors. Environmental leakage is relatively high.
Literate programming suffers from this in general because you don't want to clutter up your document with noise about versions and so forth - even as appendices, including package information and build instructions is a lot of noise.
Requiring that papers making claims about some code base publish that code base simultaneously with the paper via the peer review process should be sufficient to improve the overall state of affairs.
Don't underestimate the value of demos and competitions with companion papers either, those tend to get a lot more notice.
- disgruntledphd2 7y agoI dunno, I get where you are coming from with respect to literate programming, but I find that it's often better to show all the versions in an org file (or whatever tool you use) and write up a report separately including the final results. In general, you'll have a lot of approaches that don't work out which are nice to have a record of, but definitely don't merit being in the final paper.
- codebje 7y agoA published code base doesn't necessarily mean a full revision history, just something others can reproducibly build and run. If your claims don't depend on specific behaviours of some body of code, you wouldn't need it - eg, an article claiming some asymptotic performance of an algorithm should describe the algorithm in the abstract s.t. the performance bound can be proven, not some specific language's implementation of the algorithm.
- dv_dt 7y agoSeems like one should be able to tag some metadata to a notebook - as little as a git url and a hash or as exotic as an IPFS link or (I hate to say it) some other blockchain info. Somehow dependencies would seem to be something needed to be specified in the submittal standards of a code-required journal
- codebje 7y agoMy experience with ML notebooks in particular has been poor. But most of the publications I read are not in the field of ML, and I tend to prefer ones that use an abstract definition of an algorithm rather than code in any case - whatever the language of the day is, it probably won't be compilable in fifteen years, but the algorithm will probably still be relevant.