4 ms·
A bit of a counterpoint: http://yosefk.com/blog/why-bad-scientific-code-beats-code-following-best-practices.html http://yosefk.com/blog/why-bad-scientific-code
by dasmoth 9y ago
A bit of a counterpoint:
http://yosefk.com/blog/why-bad-scientific-code-beats-code-following-best-practices.html http://yosefk.com/blog/why-bad-scientific-code-beats-code-fo...
- belusidaty 9y agoSo, I'm a professor and, like a lot of things, think there's truth to both sides. Computer science is something that seems underrecognized in the field I work. People clearly acknowledge its importance, but then turn around and basically ignore it when talking to potential grad students or mentoring undergrads in prep for grad school. We don't offer any kind of course like "programming for X" even though most of the students need it. My experience closely parallels the "why bad scientific code beats code following best practices." I've had comp sci students come in and what happens is they clearly understood python and java, but had difficulty understanding the problems with inheritance, and wrapping their head around other more functional languages we were using. They also were unfamiliar with the statistical/content areas, so had difficulty implementing things. I had thought it would be great to have comp sci students involved (and still do) but it didn't solve my problems like I thought--so instead of having students who understood the concepts but not the programming, now I had students who understood the programming but not the concepts. When you're dealing with really intense math and statistics, it's difficult to separate out the programming from the math. It's not like web development where you have an "insert text here" kind of approach that works often; the algorithms and the problems are really wrapped up in one another. This might all be changing with data science DL and AI and that kind of stuff infiltrating comp sci's assumed territory, but I'm not really seeing it much so far. It seems the prototypical situation in software design is some software that's team-developed for mass consumption. In science, you have the reverse often, which is software designed by small units that might be a one-off thing. These constraints put different kinds of pressures on the process, such as intense pressure on getting something to work correctly at all costs, including elegant design. Also, the unit testing thing is kind of confusing to me. Every time discussion about a new language comes up in the context of numerical/scientific computing, one of the big questions is "does it have a REPL"? It seems one of the big reasons for doing this is basically unit testing. It might not be unit testing in the formal sense that you might have at some software design companies, but anything someone complicated involves feeding each tiny separable part of the code something with known expected output, sometimes in strange, boundary-testing ways, so that seems pretty similar to me. There's also a plethora of test-case datasets out there for this very purpose. To me the bigger problem is homogenization in software in science, that is, a domain being dominated by a single piece of software. I think it leads to unrecognized errors due to lack of replication across implementations, and problems typical of monopolies (even when something is open source). There's a kind of development benefit:cost supply:demand problem that leads to dominance of single works of code that is really unhealthy for science (replicating with standardized methods is good too, but to me that's a slightly different issue).
- dasmoth 9y agoThanks for the in-depth reply, a lot here that I agree with, particularly on the homogenisation/monoculture issue (which I wish I could offer a compelling answer to -- telling people not to share their code would clearly be throwing the baby out with the bathwater). Unit test vs REPLs is an interesting one. Agree that there are similarities there (although I'd argue that your tooling needs to be pretty damn good for unit tests to offer the responsiveness that a good REPL can). For me, I think part of the difference is that a REPL session is personal and nobody sees the blind alleys, while unit tests are an enduring part of the product and something others will see, use, and potentially critique. So while they can address the same kinds of questions, I'm not too shocked that people feel differently about them.