5 ms·
I think what you are describing is pretty much the only promising direction for solving the problem. However, I have significant doubts about the approach. Pe
by socratic 15y ago
I think what you are describing is pretty much the only promising direction for solving the problem. However, I have significant doubts about the approach. Perhaps you can address them?
The strategy as I understand it is:
1. Convince a high profile conference in ${FIELD} that reproducibility is important.
2. Create a special group within that community to test submitted code to see if it matches the results presented in papers.
3. Give a special carrot to authors (a special mention in the program, a piece of text in their paper) who meet the expectations of this group.
4. Hopefully, eventually readers come to see papers with the markers indicating reproducibility as the only legitimate ones, and writers are then required to make the significant time commitment (and take the significant risks) of releasing their code.
As it happens, (1), (2), and (3) have happened in a few systems communities. For example, SIGMOD has (more or less) the same setup as you describe.
However, I have deep doubts about whether (4) will ever happen. The three issues are:
1. The group doing the evaluation of the code for the conference has a boring, unappreciated job. They are also reading terrible, likely buggy code. A natural outcome is that the evaluation group will make bold claims about how all of the code they evaluated had significant issues potentially impacting research results, making everyone who submitted look bad, and leading to disincentives for future submitters. In fact, the evaluation group may even write papers about how bad specific code they reviewed was. I believe this has happened in other communities.
2. I briefly alluded to this in my original post, but many actors have extremely good reasons (at least on their face) for not releasing their code and/or data. This is why I mentioned how researchers embedded at companies modifying large proprietary code bases are extremely unlikely to ever be part of this evaluation regime. (And, no one wants to kick such researchers out of the academic community.)
3. In order for a stigma to be attached to non-reproducibility according to the conference, there has to be a strong correlation between the highest quality work and reproducibility. However, it is likely that much of the highest quality work will not be reproducible, either because it comes out of (or in conjunction with) corporate research labs, or because it uses some very difficult to get proprietary data. Likewise, the most easily reproducible results may be the least significant.
Do you think that these issues are solvable in the long term?
- stonemetal 15y agoHow is there not a stigma attached to non-reproducibility? If it is not reproducible, then the findings are incorrect, and an investigation should be made to determine if it is just incompetence or out right fraud. Otherwise fire up the nobel prize committee for I have just discovered cold fusion and cured cancer.
- St-Clock 15y agoAbout Problem 1, I must stay vague, but the code I read was understandable. We tracked the time it took us to review the artifacts and since this is still an open problem (how to efficiently and fairly review artifacts), the process may change in the future. To address the potential negative impact on the authors, I believe the papers who did not get a good review by the artifact evaluation committee just did not get a special mention. Since it is not possible to know who submitted an artifact and who did not (unless you got a mention), no harm is done for now. The potentially negative impact on the authors' reputation was something that concerned me because I've been burned in the past by Ph.D. students not being able to use the tools I published and saying that my tools were buggy when they just did not know how to install Eclipse... About problem 2, the conference organizers promised to keep the data confidential, but that might not be enough in some cases. For example, I would never show my interview transcripts to anyone, but there are some intermediate data that I could show and describe. We did not need to reproduce everything, we just wanted to see reasonable evidence that the approach described in the paper had been validated as advertised. I'm not sure I understand the difference between problem 2 and 3. I must say that in my research area, I don't see many approaches and studies that are exclusively about proprietary data. Often, some part of the data/technique is publicly available or the approach has been tried on both proprietary data and open source data. Overall, I think the artifact evaluation committee is a nice initiative and a step in the right direction. It needs to be carefully monitored and adapted to ensure that nobody gets burned for a bad reason though.
- _delirium 15y agoAs far as examples for 2/3, Google is a common source. Lots of their papers report on experiments conducted with: 1) massive-scale proprietary data; and 2) using proprietary infrastructure. They do sometimes share data, but often don't, and in cases where they don't, there is not always equivalent publicly available data (especially of anywhere near the same size). And I would suspect that sharing the code is right out for a lot of cases; if they're writing a paper on improving an aspect of Google Translate (which they do fairly regularly), they aren't going to send the entire source code to GT to a committee.