5 ms·
For CS, reproducing should be easy given the code and input data right? Other sciences' input isn't so easily shared
by studentrob 10y ago
For CS, reproducing should be easy given the code and input data right? Other sciences' input isn't so easily shared
- timwaagh 10y agomost cs papers i read (actually all) do -not- include code (or input data). which makes this an issue even in the case of cs.
- mannykannot 10y agoRerunning code does not count as reproduction as it is not an independent test of what is being claimed. You do have the option of verifying the code against the goals of the study, though that is a review rather than a reproduction (which is still valuable.)
- maze-le 10y agoNot necessarily. There are fields like complexity and computability wich are pureley mathematical in nature and have little to do with actual code, but much with the abstract theoretical fundamentals of code in general. And on the other hand you have fields like machine learning, where you can have an algorithm, wich you can implement (or have as code), but you don't nessessarily know the values of certain parameters, specific for your problem space.
- dagw 10y agoIdeally you should re-implement the algorithm based on the description in the paper to verify that the description of the algorithm is correct. You should also test with your own data to make sure that the algorithm works on all reasonable data and not only on some provided cherry picked data. If you can't get the expected results with your own implementation and your own data then the results aren't reproduced.
- studentrob 10y agoGreat points, thanks!
- eru 10y agoYes. So being able to rerun with the same code and same inputs to get the same outputs is a lower bar. Many papers don't meet even that bar. (Mostly because they don't publish code nor data; and academic code is often a horrible mess, and the code was mucked around with between different stages of running.)
- adrianN 10y agoGood luck getting the code. Abandon all hope if you want to make it run on a machine different from the laptop on which the original grad student implemented it.
- sidarape 10y agoI mean, if a paper gives an algorithm, proves its correctness, and you are convinced by the proofs, then you're done. I don't see how implementing the algorithm gives you more insight. I'm doing my master degree in computational geometry and most people in my lab don't even implement their algorithms. They just know they are correct from their proofs.
- teddyh 10y ago“Beware of bugs in the above code; I have only proved it correct, not tried it.” — Donald Knuth (https://staff.fnwi.uva.nl/p.vanemdeboas/knuthnote.pdf https://staff.fnwi.uva.nl/p.vanemdeboas/knuthnote.pdf)
- tnecniv 10y agoBecause evaluating algorithms is not always that straight forward. For some algorithms runtime is hugely important, and I don't mean the asymptotic complexity but a hard benchmark of how much it can do in what time span. Stuff like a SLAM algorithm being O(n^2) is nice and all, but to compare it to other SLAM algorithms, I need hard numbers on what it can do in how many milliseconds. Often, I find that authors don't publish their code. If they do publish their code, they rarely publish their code for their benchmarks.
- sidarape 10y agoYeah sure but at that point, it's more optimization than algorithm design. Of course, with any algorithm, you always have the hidden constant that you must account for. Also, what I was saying does not apply to the entier CS field. It only applies when you try to design an algorithm for a problem that does not yet have an efficient algorithm. I don't have much experience but I don't think it is really hard to see the cost of the hidden constant in most algorithms.