6 ms·
Research code is generally an abomination, scraped together by MS/PhD students who may never have had any industry exposure to best practices, norms, testing, e
by mayank 7y ago
Research code is generally an abomination, scraped together by MS/PhD students who may never have had any industry exposure to best practices, norms, testing, etc. There may also be the fear of releasing buggy code, which if/when discovered, will discredit the paper. Far safer to just not release the code, since it isn't mandated (which should absolutely 100% change)
Source: am former producer/publisher of abomination research code.
- dv_dt 7y agoI wonder if a journal could be established that emphasizes code integrated papers, something that might accept submissions as jupyter notebooks for example. Edit: Perhaps one way to encourage this is via topic "reuse" papers, which present further detail on a previously published papers topic, but this time with code and more detailed discussion. It would at least help the situations in that the author gets some reuse of effort, but in a way that still shows new information and advances the field.
- scribu 7y agoFormatting articles as jupyter notebooks might help with reproducibility, but it wouldn't necessarily help with understanding. distill.pub has a journal that strongly encourages interactive articles: https://distill.pub/journal/ https://distill.pub/journal/ It's a lot of work, tho: > One good test for whether your article is a fit for Distill is whether your collaborators and you are willing to put in whatever time is necessary to write and illustrate an outstanding article. In our experience, this often takes 100+ hours.
- dv_dt 7y agoReproducibility is one of the core problems of the article. I'm not sure any format can assure understanding by itself - that's more the purview of the reviewers I would think. Applying the criticism to my own suggestion: even the reproducibility issue is only tangentially helped by a Jupyter format, it may only bring reviewers to the point where one might execute the notebook (though there are still issues there if its an exceeding long computation or needs special hardware setup, or the notebook executes w/o demonstrating actual computation).
- tsumnia 7y agoThis is a current area of research as the prevalence of interactive notebooks has peaked. As of now, they simply "EXIST" and we know we like them. Its still unclear as to why; obviously having code intertwined other common communication mediums but what about that makes us like it. What benefits does it have over pseudocode or just code? What about code beyond small scripts (500+ lines with design patterns or other abstractions)?
- codebje 7y agoJupyter notebooks are risky business for reproducible work. No dependency data means they're highly prone to bit rot. Stored results and out-of-order execution means they're prone to subtle errors. Environmental leakage is relatively high. Literate programming suffers from this in general because you don't want to clutter up your document with noise about versions and so forth - even as appendices, including package information and build instructions is a lot of noise. Requiring that papers making claims about some code base publish that code base simultaneously with the paper via the peer review process should be sufficient to improve the overall state of affairs. Don't underestimate the value of demos and competitions with companion papers either, those tend to get a lot more notice.
- disgruntledphd2 7y agoI dunno, I get where you are coming from with respect to literate programming, but I find that it's often better to show all the versions in an org file (or whatever tool you use) and write up a report separately including the final results. In general, you'll have a lot of approaches that don't work out which are nice to have a record of, but definitely don't merit being in the final paper.
- codebje 7y agoA published code base doesn't necessarily mean a full revision history, just something others can reproducibly build and run. If your claims don't depend on specific behaviours of some body of code, you wouldn't need it - eg, an article claiming some asymptotic performance of an algorithm should describe the algorithm in the abstract s.t. the performance bound can be proven, not some specific language's implementation of the algorithm.
- dv_dt 7y agoSeems like one should be able to tag some metadata to a notebook - as little as a git url and a hash or as exotic as an IPFS link or (I hate to say it) some other blockchain info. Somehow dependencies would seem to be something needed to be specified in the submittal standards of a code-required journal
- dilawar 7y agoThere are some journals who have started taking it seriously. eLife does that: asks authors to put their code on GitHub/gitlab etc. Fork their repo and use that as snapshot for the published paper. I like the eLife approach. But for most journals, it's the story as usualy. Personally I dont feel very confident about a paper result if they don't publish the code. Why hold back on the code?
- khawkins 7y agoI disagree that it is a lack of industry exposure. "Best practices, norms, testing, etc." is a waste of time when you develop well written code that doesn't produce good results. After months of tweaking even a good foundation of code, it begins looking atrocious. But you get graduated for publications based on good results, not for code that someone else can use.
- kian 7y agotesting, refactoring, code organization, consistent code styling for easy search-replace, and other best practice norms will save you months of effort in tracking down bugs, fixing new ones you introduced while scrambling around researching, figuring out what's going on when you have an idea that uses code you wrote six months ago, etc. Sure, maybe a tutorial and documentation might not be worth the effort for something that hasn't produced results (does publishable research still count as not having produced results?), but most industry best practices I can think of will save you much more than the time invested even on relatively modest (< 1 month) projects.
- majormajor 7y ago> After months of tweaking even a good foundation of code, it begins looking atrocious. Months of tweaking code is exactly what you want some test coverage for - so that you end up with a tweaked version of something where you still know what it does instead of something that does god knows what.
- closeparen 7y agoTyping “git commit” when you reach a milestone costs nothing, and can save hours of trying to get back to a previous working state.
- Reelin 7y ago> Research code is generally an abomination ... http://matt.might.net/articles/crapl http://matt.might.net/articles/crapl > The CRAPL is an open source "license" for academics that encourages code-sharing, regardless of how much how much Red Bull and coffee went into its production.
- Someone 7y agoIn addition, you may want to milk your code for more papers. Releasing the code makes it easier for others to also do that milking, too, finding and publishing some results before you can find them.
- solveit 7y ago> There may also be the fear of releasing buggy code, which if/when discovered, will discredit the paper. I know you know this, but this is absolutely a feature.
- hn23 7y agoWell, he was asking for pseudo code. So this could be at a high level where you do not really have to fear bugs, do you?
- skunkpocalypse 7y agoYeah the big question isn't why code isn't published. The big question is why this failure to publish code is tolerated by program+steering committees.