32 ms·
If you have a good paper with important result, providing code and data is not necessary. Providing the code to replicate is good form. It shows good faith and
by MAXPOOL 8y ago
If you have a good paper with important result, providing code and data is not necessary.
Providing the code to replicate is good form. It shows good faith and confidence. Exact replication (exactly replicating the study) is just the starting point to check that the code works and no obvious mistakes were made.
replication / reproducibility / hyperparameter sensitivity
If the research yields something really important and the method is well documented, usually it can be easily checked without having the data and the code. Things like dropout, batch normalization, residual learning, .. work over multiple different datasets and hyperparameters. You can reproduce the results without faithfully replicating the experiment.
If the claimed result vanishes unless you have the exact data, code or the hyperparameters, the research can't be said to be meaningfully reproducible in the scientific sense. Hyperparameter sensitivity is ML equivalent to P-Hacking.
- lumost 8y agohow many papers important results are simply bugs? Numerical code is already bug prone due to subtle and hard to test errors, research code that's not code reviewed or necessarily even tested for correctness can easily generate important results erroneously. It's also counter-productive not to publish the underlying source code for these papers, as it adds a barrier to other researchers applying the algorithm in new situations. I'd be interested in seeing if those 6% of papers which include the code get more citations than the population of papers which do not include code.
- MAXPOOL 8y ago> important results are simply bugs Probably none. If the paper is important and collects citations, the algorithm is in use. Computer science != working code. Code is required when you produce something where the scientific importance is less clear. There is need to provide more evidence. Many papers are just "Hey I made some some tweaks and it works in this particular case." Those papers should have working code.
- lumost 8y agoMost papers leverage a results table which compares the newly proposed approach with existing approaches. This section is baselines the results of the new approach with prior work and helps determine whether a new result is actually important or yet another way to achieve the same results as previous work. e.g. the tables on page 7 of this paper https://www.semanticscholar.org/paper/Automatic-Acquisition-of-Lexical-Formality-Brooke-Wang/823a397b29bf596e2734d3dff7ab5abec2f60ac9 https://www.semanticscholar.org/paper/Automatic-Acquisition-... These tables are generated using real implementations that may or may not be correct, and should be subject to review when the paper is published.
- alevskaya 8y agoAs someone who's often implementing models from ML papers, the issue is that english-language descriptions of methods is often found to be sorely lacking. Even good authors simply forget to mention small critical details of model design or training that wouldn't remain ambiguous if they simply released their code. Part of the effort of reproducing work is certainly figuring out which aspects of the model design are critical or incidental - but this is all greatly aided by not having to guess at what was actually done from a heap of english and LaTeX generally written feverishly over a few days before a deadline!
- iguy 8y ago> Part of the effort of reproducing work is certainly figuring out which aspects of the model design are critical or incidental Isn't this the work of writing up? A paper is a claim that you have discovered something, and implicitly that other details are standard / unimportant. If it turns out that some hidden assumption was in fact doing all the work, then the claim you made was false. I'm all in favor of sharing working code. But working code which magically does something... amounts to an anomaly awaiting an explanation. Or an advertisement.
- sdenton4 8y agoAlas, training times are quite long... At any given time, there's probably a couple best-in-class architectures out there for the problem you're interested in, and one or two dozen interesting bells and whistles one can add as decoration, each with a paper that makes pretty reasonable arguments and has some stats demonstrating modest gains. The right thing to do in this circumstance is an ablation study - throw together your best-possible model and then test different subsets of features sitting between your model and the 'basic' prior work. For large datasets, though, each of these models might take a very long time to train (especially if you don't work at a place with a stupid number of GPUs available). So, lacking resources, you get your new best-ever accuracy number with your 'everything' model, and do an extensive write-up about how awesome the new bell and/or whistle that you added to the pile is... (The problem is compounded by a need to publish quick, lest someone else describe your bell/whistle first.) Another big problem is that adding a bell/whistle to the base model often means adding more parameters to the model. There's decent evidence coming out of the AutoML world that number of parameters matters a hell of a lot more than how you arrange them. (It's real real easy to convince yourself that your clever new idea is more important than the shitpile of new parameters you've added to the model, after all.) So a really solid ablative study probably needs to scale the number of parameters in a reasonable way as you add/remove features... And there may not be obvious ways to do that smoothly. And this is closely related to the replication study in psych: it's real real easy to do kinda sloppy work with big words attached, and convince yourself and all your peers that you're a genius. I think a big database for reporting and searching for results with various architecture+dataset combinations would be much more useful than pushing more papers to the arxiv in many cases. (though, really, whynotboth.gif) Let me do some searches to see if a particular bell/whistle actually adds value across the god-knows-how-many-times someone's used it to train up imagenet from scratch...
- abecedarius 8y agoReproducible builds make it much easier for others to check for things like hyperparameter sensitivity. Right?
- MAXPOOL 8y agoYes.