4 ms·
Unfortunately, Jacques needs to answer a pretty basic question in order to get his wish: why? Why should CS paper writers show Jacques (or anyone else) their c
by socratic 15y ago
Unfortunately, Jacques needs to answer a pretty basic question in order to get his wish: why? Why should CS paper writers show Jacques (or anyone else) their code?
We have discussed this topic on HN a number of times, for example:
http://news.ycombinator.com/item?id=2735537 http://news.ycombinator.com/item?id=2735537
http://news.ycombinator.com/item?id=2006749 http://news.ycombinator.com/item?id=2006749
Many of the comments in those threads do a better job summing up than I ever could. However, briefly, literally all of the incentives are aligned against publishing code and data.
If a writer's code is wrong, they are embarrassed (and there is no culture of being embarrassed by not publishing code).
If a writer publishes their code and it is actually good, someone else can scoop their follow-on results.
If a writer does not publish their code, and it is actually any good, they can potentially commercialize it thanks to the Bayh-Dole Act.
If a writer publishes their code and people intend to use it, the writer needs to clean it up, check it for correctness, and handle support requests. These activities are probably more time consuming than writing the code in the first place.
If the writer publishes their code, and other people in the writer's field do not, the writer is usually at a disadvantage. Others will appear to have more publications, the basic currency of academia. (Many people have great reasons for not publishing their code or data, especially researchers embedded at large companies making changes to large proprietary systems.)
So overall, yes, it would be great if CS paper writers gave out their code. What they are doing is not reproducible science in the philosophy of science sense.
But what is Jacques (or anyone else) doing to fix this system of incentives, and what could anyone do?
- bartl 15y agoIt is all for science. Other people can validate you didn't make any blatant errors.
- felipemnoa 15y agoFor validation purposes is probably better that the researches doing the validation write their own code from scratch.
- _delirium 15y agoIndeed, that's the main component of the gold-standard results replication. For a chemistry experiment, for example, you're not supposed to replicate it by going to the original lab, using their existing apparatus, and just re-running the experiment. Instead, there's stronger confidence in the results if you replicate it using your own equipment in your own lab, reconstructing any necessary components from descriptions in the paper. That way you know that the results were actually due to what the paper claimed they were, instead of some overlooked happenstance in the original lab or apparatus. Of course, reimplementation can be quite time consuming, which is the main problem. But then sharing code can actually decrease the likelihood of anyone ever reimplementing the algorithm again, instead just re-using the same (possibly buggy) code forever without looking at it.
- wisty 15y agoOK, but there's no way subtle bugs will be found unless the code is released. If you replicate a non-trivial piece of code, both will have bugs, and both will have subtle design trade-offs (which won't be documented in the paper). Will you publish your results, despite the fact that they don't agree with the existing, accepted ones? If you do, what will it achieve?
- skrebbel 15y agoIn science, there are still such things are "good manners" and "common scientific practices", not all of which are directly incentivised. For example, every article contains references to other articles that helped create it. Every article credits all authors, usually in a particular order (which differs culturally per discipline). They're all fairly common sense, and there are many more. Some are listed explicitly as rules for journal or conference submissions. Some, however, are considered too obvious for that. These are rules everyone follows. People get frowned upon when they are not followed. Unfortunately, when these rules came to be, there wasn't that much code around. You can't print out your physics experiment setup. Additionally, every additional piece of info, be it data, code, or whatever, would need to be printed and shipped. This cost too much, which encouraged a culture in which only the most important parts, findings and a brief story of how they were obtained, were to be included in publications. But then the internet came. The whole world adapted, but science culture didn't. The internet allows large sets of information to be shipped along with papers. This does not just hold for computer science. Entire data sets of physics experiments could be included. Survey data, code, Matlab models, and so on. All this data is collected using tax payers' money, done for the greater good of the world. It makes perfect sense to include it all with the publication, online. In fact, it is criminally wasteful to make different teams collect the same data, or do the same work, over and over again. I think the only change necessary to make this happen is for a few influential publications and conferences to require that all data, code, etc is published online along with the paper. It can start in a niche subject area and expand from there. What Jacques (or anyone else) is doing to fix this system is yelling on the internet about it, so that conference programme committees and journal reviewers may at some point be swayed to add this little rule. A rule that makes sense, just like "your paper must contain an abstract" makes sense, just like "you must include references to all works you borrowed from or built upon" makes sense.
- Hyena 15y agoHe's not actually saying they lack incentive so much as that there are significant disincentives. The rest of sience contains an incentive, various types of professional censure, for the behavior you describe. CS does not and in fact contains significant ince tives for secrecy.
- St-Clock 15y agoThanks socratic for summarizing well the previous discussions! I was quite vocal against the obligation to release code if there was no incentive, even though I released the code of all the papers I published so far [1]. But to my surprise, the software engineering research community decided to try something new this year in one of their top conferences, FSE [2]. All authors of accepted papers were invited to submit their artifacts (data, code, video showing how the code was used, etc.) to the artifact evaluation committee. Two members of the committee reviewed each artifact and compared it with the paper. Papers that received a grade equal or above "meet expectations" received a special mention in the program [3]. Although the artifacts are not released to the public, this is a step in the right direction. If you are motivated to package your code for artifact evaluation and you get some recognition out of it, then the next step of releasing it is a lot easier. As a committee (I was part of it), we were afraid it would take hours and hours to just try to run the code that was submitted, but the authors really made the effort to package their code well. [1] http://news.ycombinator.com/item?id=1868581 http://news.ycombinator.com/item?id=1868581 [2] http://2011.esec-fse.org/cfp-artifact-evaluation http://2011.esec-fse.org/cfp-artifact-evaluation [3] http://2011.esec-fse.org/program-details http://2011.esec-fse.org/program-details
- socratic 15y agoI think what you are describing is pretty much the only promising direction for solving the problem. However, I have significant doubts about the approach. Perhaps you can address them? The strategy as I understand it is: 1. Convince a high profile conference in ${FIELD} that reproducibility is important. 2. Create a special group within that community to test submitted code to see if it matches the results presented in papers. 3. Give a special carrot to authors (a special mention in the program, a piece of text in their paper) who meet the expectations of this group. 4. Hopefully, eventually readers come to see papers with the markers indicating reproducibility as the only legitimate ones, and writers are then required to make the significant time commitment (and take the significant risks) of releasing their code. As it happens, (1), (2), and (3) have happened in a few systems communities. For example, SIGMOD has (more or less) the same setup as you describe. However, I have deep doubts about whether (4) will ever happen. The three issues are: 1. The group doing the evaluation of the code for the conference has a boring, unappreciated job. They are also reading terrible, likely buggy code. A natural outcome is that the evaluation group will make bold claims about how all of the code they evaluated had significant issues potentially impacting research results, making everyone who submitted look bad, and leading to disincentives for future submitters. In fact, the evaluation group may even write papers about how bad specific code they reviewed was. I believe this has happened in other communities. 2. I briefly alluded to this in my original post, but many actors have extremely good reasons (at least on their face) for not releasing their code and/or data. This is why I mentioned how researchers embedded at companies modifying large proprietary code bases are extremely unlikely to ever be part of this evaluation regime. (And, no one wants to kick such researchers out of the academic community.) 3. In order for a stigma to be attached to non-reproducibility according to the conference, there has to be a strong correlation between the highest quality work and reproducibility. However, it is likely that much of the highest quality work will not be reproducible, either because it comes out of (or in conjunction with) corporate research labs, or because it uses some very difficult to get proprietary data. Likewise, the most easily reproducible results may be the least significant. Do you think that these issues are solvable in the long term?
- shabble 15y agoPart of the problem is that academic progression is based so heavily on churning out papers that there's a race to the bottom for 'minimum publishable units'. A single codebase might provide the analysis framework for tens of individual papers, and releasing early just increases the chances of getting scooped on some of them. I think there are precedents for delayed data distribution in other fields (biology / chemistry?), so maybe a partial solution would be to embargo the code for a year or so after publication. The downside is that you can never expect to see code for the really cutting edge stuff, but perhaps that will be the price to pay? Having a delay also might serve to blunt the "Yeah, well. Your code sucks." argument, in that it would need to really suck in a massive, conclusion changing way, in order to gather new headlines. Otherwise, it just sort of sinks away into obscurity unless you're looking directly at that result, in which case knowing about possible defects and improvements is a net benefit. Overall, I'm offering no new benefits here, but at least a couple of possible mitigations of the downsides you mention.
- scott_s 15y agoWhen I was a grad student, I put all of my code online for my papers. I did this for two reasons: basic transparency, and the hope that it would be useful to others. I actually got lucky a few times, and people did think my code was useful. It was used in other research projects, which got me citations - which are academic incentives. I have bunted on a few support requests, though. Regarding the disincentive that publishing code means others may scoop your own work, I find that doubtful. By the time a conference paper is published, almost six months have passed since the paper submission. Most researchers are well on their way to writing another paper with that work and codebase by the time the wider community knows about it. Then there's the difficulty of making non-trivial changes to prototype-quality code implementing something new.