7 ms·
What possible reason is there for not doing so? The code and the ideas within have no direct commercial value, surely? Laziness, lack of time, and the fact tha
by hugh3 15y ago
What possible reason is there for not doing so? The code and the ideas within have no direct commercial value, surely?
Laziness, lack of time, and the fact that most scientific code is not suitable for public consumption. For instance, my code includes error messages such as "What the fuck????" and "This never happens" which I'd need to take out in order to prepare it for publication.
In addition, there's the assumption that source code is a pretty trivial implementation detail; we publish the algorithm but not the details of the implementation, just as the experimentalist tells us what he built but doesn't tell us what brand of screwdriver he used to build it.
- JonnieCache 15y agoThe thing is though, as I'm sure you're aware, the algorithm in abstract is utterly irrelevant. The algorithm as released is never actually executed, it is the implementation that gets executed. One "trivial" detail in the implementation, one tiny little floating point rounding error, can throw the entire thing. We all know this. So why is the scientific community so unwilling to face this elephant in the room? And why is nobody else confronting them about it? Of course there's jgc, with his admirable record of confronting people about things, but that's the only example I can think of. You would've thought the "Al Gore made up AGW to devalue my exxon shares" angst-brigade would've jumped on this, but they seem more concerned with digging through departmental email gossip.
- jonhendry 15y ago"The algorithm as released is never actually executed, it is the implementation that gets executed." If you use a different implementation of the algorithm, on the same data, and get a different result, you will find that there was an error. That's a lot easier than combing through someone else's code base. A simple example: If Lab A says they took the square root of 16, you don't need their source code to know there's a problem, if their result is 5 and your result is 4. If you use their code, you might not notice that it's borked.
- rbanffy 15y agoStill, this is not a good reason not to release code. We can deal with funny comments, bad jokes and dead code. Why not let people examine the code? They are not programmers who will be judged by the neatness of their code.
- jonhendry 15y agoWhat if they don't have code, because they're using a commercial product to do the recording?
- rbanffy 15y agoWell... That part has to be certified as a black box, but even if your code is a Mathematica notebook, it would be useful to have it published along the raw data. Again, there is no reason to be embarrassed of comments like "this should never happen". My own code has parts that raise "ThisShouldNeverHappenError" when something that should never happen disregards my opinion and happens anyway. ;-)
- jerf 15y agoYou've implicitly accepted that the first implementation is correct, which I've found is one of the mental blocks that people defending this practice seem to have with frequency that surprises me. It isn't true. It isn't even the sort of thing you can try to debate your way out of, because it simply isn't true. Rather complicated numeric code written by non-software engineers is not something I'm willing to presume correct. I wouldn't presume it correct if it were written by software engineers, either, but at least then we'd probably have some sort of halfway reasonable testing to point at. Further, the Kolmogorov complexity of a simulation is typically a wee bit larger than the Kolmogorov complexity of a square root implementation. That metaphor is beyond useless, it's deceptive.
- jonhendry 15y ago"You've implicitly accepted that the first implementation is correct" Not at all. The questions are 1) does the method described sound reasonable and 2) if I implement the same methods, with my toolset, and process the same data, do I get the same result? If the results are different, using different implementations of the same methods, then it's possible that either party could have the error. But if the second party is reusing well-tested code of their own, or something from Matlab or whatever, then it is more likely that the unknown code of the original lab is the problem. You might have a point about a simulation, but an awful lot of science isn't simulations, it's analyzing recorded data.