11 ms·
Thanks socratic for summarizing well the previous discussions! I was quite vocal against the obligation to release code if there was no incentive, even though I
by St-Clock 15y ago
Thanks socratic for summarizing well the previous discussions! I was quite vocal against the obligation to release code if there was no incentive, even though I released the code of all the papers I published so far [1].
But to my surprise, the software engineering research community decided to try something new this year in one of their top conferences, FSE [2]. All authors of accepted papers were invited to submit their artifacts (data, code, video showing how the code was used, etc.) to the artifact evaluation committee. Two members of the committee reviewed each artifact and compared it with the paper. Papers that received a grade equal or above "meet expectations" received a special mention in the program [3].
Although the artifacts are not released to the public, this is a step in the right direction. If you are motivated to package your code for artifact evaluation and you get some recognition out of it, then the next step of releasing it is a lot easier. As a committee (I was part of it), we were afraid it would take hours and hours to just try to run the code that was submitted, but the authors really made the effort to package their code well.
[1] http://news.ycombinator.com/item?id=1868581 http://news.ycombinator.com/item?id=1868581
[2] http://2011.esec-fse.org/cfp-artifact-evaluation http://2011.esec-fse.org/cfp-artifact-evaluation
[3] http://2011.esec-fse.org/program-details http://2011.esec-fse.org/program-details
- socratic 15y agoI think what you are describing is pretty much the only promising direction for solving the problem. However, I have significant doubts about the approach. Perhaps you can address them? The strategy as I understand it is: 1. Convince a high profile conference in ${FIELD} that reproducibility is important. 2. Create a special group within that community to test submitted code to see if it matches the results presented in papers. 3. Give a special carrot to authors (a special mention in the program, a piece of text in their paper) who meet the expectations of this group. 4. Hopefully, eventually readers come to see papers with the markers indicating reproducibility as the only legitimate ones, and writers are then required to make the significant time commitment (and take the significant risks) of releasing their code. As it happens, (1), (2), and (3) have happened in a few systems communities. For example, SIGMOD has (more or less) the same setup as you describe. However, I have deep doubts about whether (4) will ever happen. The three issues are: 1. The group doing the evaluation of the code for the conference has a boring, unappreciated job. They are also reading terrible, likely buggy code. A natural outcome is that the evaluation group will make bold claims about how all of the code they evaluated had significant issues potentially impacting research results, making everyone who submitted look bad, and leading to disincentives for future submitters. In fact, the evaluation group may even write papers about how bad specific code they reviewed was. I believe this has happened in other communities. 2. I briefly alluded to this in my original post, but many actors have extremely good reasons (at least on their face) for not releasing their code and/or data. This is why I mentioned how researchers embedded at companies modifying large proprietary code bases are extremely unlikely to ever be part of this evaluation regime. (And, no one wants to kick such researchers out of the academic community.) 3. In order for a stigma to be attached to non-reproducibility according to the conference, there has to be a strong correlation between the highest quality work and reproducibility. However, it is likely that much of the highest quality work will not be reproducible, either because it comes out of (or in conjunction with) corporate research labs, or because it uses some very difficult to get proprietary data. Likewise, the most easily reproducible results may be the least significant. Do you think that these issues are solvable in the long term?
- stonemetal 15y agoHow is there not a stigma attached to non-reproducibility? If it is not reproducible, then the findings are incorrect, and an investigation should be made to determine if it is just incompetence or out right fraud. Otherwise fire up the nobel prize committee for I have just discovered cold fusion and cured cancer.
- St-Clock 15y agoAbout Problem 1, I must stay vague, but the code I read was understandable. We tracked the time it took us to review the artifacts and since this is still an open problem (how to efficiently and fairly review artifacts), the process may change in the future. To address the potential negative impact on the authors, I believe the papers who did not get a good review by the artifact evaluation committee just did not get a special mention. Since it is not possible to know who submitted an artifact and who did not (unless you got a mention), no harm is done for now. The potentially negative impact on the authors' reputation was something that concerned me because I've been burned in the past by Ph.D. students not being able to use the tools I published and saying that my tools were buggy when they just did not know how to install Eclipse... About problem 2, the conference organizers promised to keep the data confidential, but that might not be enough in some cases. For example, I would never show my interview transcripts to anyone, but there are some intermediate data that I could show and describe. We did not need to reproduce everything, we just wanted to see reasonable evidence that the approach described in the paper had been validated as advertised. I'm not sure I understand the difference between problem 2 and 3. I must say that in my research area, I don't see many approaches and studies that are exclusively about proprietary data. Often, some part of the data/technique is publicly available or the approach has been tried on both proprietary data and open source data. Overall, I think the artifact evaluation committee is a nice initiative and a step in the right direction. It needs to be carefully monitored and adapted to ensure that nobody gets burned for a bad reason though.
- _delirium 15y agoAs far as examples for 2/3, Google is a common source. Lots of their papers report on experiments conducted with: 1) massive-scale proprietary data; and 2) using proprietary infrastructure. They do sometimes share data, but often don't, and in cases where they don't, there is not always equivalent publicly available data (especially of anywhere near the same size). And I would suspect that sharing the code is right out for a lot of cases; if they're writing a paper on improving an aspect of Google Translate (which they do fairly regularly), they aren't going to send the entire source code to GT to a committee.