4 ms·
Forgive the naive question: but is code like this typical for "data science"? It seems more like something out of a masters project... Huge script files with z
by physPop 5y ago
Forgive the naive question: but is code like this typical for "data science"? It seems more like something out of a masters project... Huge script files with zero testing?
Theres a huge amount of data reshaping, mutating, and general wrangling going on here, how can one be confident without even a known input->output integration type test?
- ethanbond 5y agoIn my experience, yes. Data science is highly creative/improvisational and the resulting artifacts of work reflect that. FWIW, normal science is pretty similar. The clear step 1, step 2, ..., you see in scientific publications is effectively retcon done for the benefit of the reader's understanding.
- nerdponx 5y agoYes, sadly. Testing is hard. The researchers who write this code typically don't have the necessary hands-on experience to write good tests, even if they had enough time in a day/week to actually do it. Edit: Also the code tends to be very "high level", chaining lots of high-level API functions together, and even coming up with assertions to test tends to be a bit of a challenge. Testing such code turns out to be surprisingly difficult; you might end up just rewriting big chunks of your code in the test suite. In my data science work, I've focused on writing tests for the complicated sections (e.g. lower-level string processing routines) and just trying to focus on the keep the other stuff very clean and readable.
- StavrosK 5y agoWhat would "good tests" look like for something like this?
- physPop 5y agoGenerate synthetic data -> run model -> check expected outputs. Yes it a lot of work, but you're reaching millions of people with this model and correctness is paramount! Similarly, even such simple test harnesses help when yourself or other go to modify the code. Having flags for "Did I break something" is very important.
- StavrosK 5y agoThe deliverable here isn't the code, it's the output. If "check expected outputs" is a subjective step anyway, why not just check that the output of the code when run on the data is as expected and skip the test altogether?
- stonemetal12 5y ago> If "check expected outputs" is a subjective step anyway It isn't. It is the evidence that calculations in the code were done correctly. It is like the problems in high school math class. The teacher controls the inputs so that the outputs are known. The teacher then runs the test problem past the student (aka the code). If the known correct answer isn't generated then we know the student didn't do it right.
- StavrosK 5y agoHow do you tell that the calculations were done correctly? Presumably you have some way of doing that. Then, why don't you apply that way to the output? You don't need tests if you only need to do it once.
- pigeonhole123 5y agoChecking that the code outputs what you think it does is a pretty low bar to clear. The fact that you can think of higher bars doesn't invalidate the value of checking this.
- stonemetal12 5y ago>How do you tell that the calculations were done correctly? Presumably you have some way of doing that. Isn't that what we are discussing, how to tell if a piece of code works? You could do a formal proof that the code works (which is long and tedious) or you can test it (less long and tedious but less rigorous). >You don't need tests if you only need to do it once. It doesn't matter how many times you plan on running it. What counts is how much we value the output. I wouldn't bet $20 that untested code works correctly.
- nerdponx 5y agoThat's part of the problem. In general (and in my opinion), a "good test" is one that asserts that an invariant is always so, or that a property that is expected to hold under certain conditions does indeed hold under those conditions. Defining such invariants and properties for "data science code" tends to be difficult, and even when you define them it might be difficult or impossible to test them in a straightforward fashion. And that's even before you get to the probabilistic stuff.
- alexpetralia 5y agoIn general I think data science code is more "research-oriented", meaning that it is not run continuously like a software app and its requirements often evolve as more information is learned from the data. Research code produces results maybe once or twice a day (manually triggered) while software apps potentially service hundreds of thousands of requests a day. Research code, because its requirements are not fixed and it is not frequently run, doesn't need to be as "stable" as a bonafide app. For a bonafide app, requirements do not change as often and the app is run virtually 24/7. Once the research code becomes "productionized" however - i.e. it is deployed in an online system where uptime & accuracy matter - then I think absolutely it becomes more engineering-heavy and looks quite a bit less like this code. Would be curious to hear others' thoughts on this distinction between research vs. production code however.
- admissionsguy 5y agoMany of the "good practices" in software development are primarily meant to reduce the cognitive effort on the part of the person reading the code. I suspect it will be a downvotably unpopular opinion, but a typical researcher has a significantly larger cognitive capacity than a typical software developer. Consequently, a piece of code of certain complexity will look simpler, and be easier to manipulate for a researcher than for an average software developer.
- patates 5y agoIt's always easy if you're the only one writing/maintaining the code. Also, if your job is writing code, I'd bet you'd have less difficulty manipulating spaghetti code.
- codyb 5y agoHuh? A researcher has greater cognitive capacity? That's a funny take. My impression was -> Software developers main focus is developing software so they spend a HUGE amount of time developing software, maintaining software, noticing patterns and bugs and pitfalls in software and thus they get pretty decent at writing software. Data scientists main focus is developing models, so they spend a HUGE amount of time developing models, tweaking models, finding data, and cleaning data, and write basic software to achieve some of those goals. I wouldn't expect Albert Einstein, Mozart, or Beyoncé to be some fantastic software developer just cause they're smart individuals. I'd expect people who spend a lot of time writing software to generally be the ones who write decent software.
- admissionsguy 5y agoThis code is above average by the standards of academia I've dealt with. You've got comments announcing what each part of the code is doing, and there is only one copy of each script. Normally, you would expect to have multiple nested directories with various variations of the same code, plus multiple variations of each script with cryptic suffixes.
- fullshark 5y agoYes, data scientists frequently are working alone and don't care about readability so much as results. Also the code isn't going into production so they don't care about optimizing it by and large.
- checker 5y agoSure, but physPop's concerns still stands - how can one be confident about the results? Is it just eyeballing "this looks right"? If so, how many anomalies are missed due to handwaving, and how many results are inaccurate? I'm genuinely curious because I understand the need to move fast but is accuracy a necessary sacrifice? (or is there a trick I don't know about)
- fullshark 5y agoThere's definitely a greater risk of a bug leading to misleading results. There's no real unique trick other than possibly someone else trying to replicate the results and catching an error, or trying to use the code on another data set and catching a mistake.
- alexpetralia 5y agoIf I understand correctly, tests are for what the code should do. There is some business logic and you are asserting that the business logic does what it should do (the test). If you manipulate the business logic, and it no longer does what the test says the code should do, the test fails. Here it is not as clear what the code should do. What should the amount of excess deaths be? In what ways would we change the logic such that the test case would break? If the input data set is static, isn't it more like a mock anyway? I think for this reason you often see more sanity checks in research code because the should case is not as clearly defined.
- stonemetal12 5y ago>What should the amount of excess deaths be? With several fake known inputs and there associated outputs we should be able to determine if the calculation is right. The result on the real world data is not known but when calculating a statistic you should be able to figure out if you are calculating the right statistic or returning 42 for all inputs.
- kspacewalk2 5y agoAs already pointed out, this isn't an engineering product. They've released the source, did some sanity checks, and moved on with their busy lives. If you find an error - cool, go ahead and report it.
- Glavnokoman 5y agoBy sanity check you mean shows the numbers I wanted to see? Not saying there's necessarily something wrong in this particular code but I know first hands how little effort is put into validating the scientific codes and how much of produces just random crap.
- kspacewalk2 5y agoWhich is why releasing their code is critical. They lay it all out for you, with their pride/reputation on the line if it produces 'just random crap'. Writing exhaustive unit tests is probably not something they are good at, nor is it a requirement for them, nor is it the best use of their time. I too know first hand that students and even academics who aren't properly trained and given the right tooling will write prototype code that's not production ready. Occasionally there's even a bona fide mistake, though usually it's just very un-generalizable. It's part of my job to make some of it more production ready. That's the best use of my time and training. Conversely, I'm shit at generating useful research ideas or writing papers. Just don't have the right combination of intuition/training/experience. Good thing my academic institution has the resources to employ them and me.
- epistasis 5y agoHow can you know which numbers you want to see? With tests it is often best practice to write the tests before writing the code that gets tested. But what if you don't know ahead of time, and can't possibly ever know ahead of time? Writing "tests" for this would be about adding stuff afterwards, like "the max and min values of this column were X and Y". But that is expected to break, if anything changes, because it's not testing anything useful. My question is: what is one concrete test here that would be useful and actually provide confidence that the code is doing what it should be doing, and how does that test provide better sanity checking than inspecting the tables at each step of the process?
- MattGaiser 5y agoFor something like this, the code is more an aid to analysis rather than the product itself.
- mistrial9 5y agoin R particularly, there is a wide range of code quality, by developer standards, yes It should be said that the statistics themselves do make for some 'guide rails' .. the stats are quite demanding, require a lot of upper-division training to use correctly, and visual feedback can course-correct as the results are iteratively found. As said in other comments, the content is often associated with a research goal that is weighted more than code-quality. In contrast, a general purpose development language has very broad application, and probably can go wrong in very broad ways. The coder is getting feedback on results, but often lots of code review and expectation of professional results. I picked out an R ETL sequence from an article last year, and I still pull it out once in a while, to see the extensive, clever and (to my eye) really hard to read R data ingestion and manipulation. Personally I think it is a fair tradeoff to say that the expertise to use this environment is a bit of a filter, and the rigor that the results (therefore intermediate results) demand, does move the expectations on code quality .. As for tests, there is probably no defense on not having tests also.. it is probably going to be more common as the field of R and data analysis inevitably grows.
- da39a3ee 5y agoYes, you can see the divide in the popularity of interactive notebooks like jupyter and observable. To programmers who use a standard git-based workflow, the interactive notebooks are anathema because you can't use version control to compare code state and monitor progress, and because of the global variables. But to many quantitative researchers for whom software version control isn't a large part of their professional worldview, it seems normal to just keep adding to the same document until you hit a state that seems to work, without much emphasis on checkpointing and navigation between checkpoints, and testing at each checkpoint.
- omginternets 5y agoThis is pretty good by data-science (and other research) standards.
- lanevorockz 5y agoThat is quite common for R language, it is a scripting language that you are suppose to explore as you use. So if you want to test it, it's very simple ... Just run the script until the part you are interested and then plot the hell out of it. It's widely known on the "theory of testing" that data oriented systems can't be tested without immense amount of effort. If you are interested there are plenty of academic talks about the subject.
- physPop 5y agoWould love some commentary on why this is an unpopular question. Good software practices don't only have to be for "production" or CI or large teams of coders. Testing for correctness could be seen as part of the work of delivering a high quality product/graph/prediction/model.