4 ms·
Sure, but physPop's concerns still stands - how can one be confident about the results? Is it just eyeballing "this looks right"? If so, how many anomalies are
by checker 5y ago
Sure, but physPop's concerns still stands - how can one be confident about the results? Is it just eyeballing "this looks right"? If so, how many anomalies are missed due to handwaving, and how many results are inaccurate?
I'm genuinely curious because I understand the need to move fast but is accuracy a necessary sacrifice? (or is there a trick I don't know about)
- fullshark 5y agoThere's definitely a greater risk of a bug leading to misleading results. There's no real unique trick other than possibly someone else trying to replicate the results and catching an error, or trying to use the code on another data set and catching a mistake.
- alexpetralia 5y agoIf I understand correctly, tests are for what the code should do. There is some business logic and you are asserting that the business logic does what it should do (the test). If you manipulate the business logic, and it no longer does what the test says the code should do, the test fails. Here it is not as clear what the code should do. What should the amount of excess deaths be? In what ways would we change the logic such that the test case would break? If the input data set is static, isn't it more like a mock anyway? I think for this reason you often see more sanity checks in research code because the should case is not as clearly defined.
- stonemetal12 5y ago>What should the amount of excess deaths be? With several fake known inputs and there associated outputs we should be able to determine if the calculation is right. The result on the real world data is not known but when calculating a statistic you should be able to figure out if you are calculating the right statistic or returning 42 for all inputs.
- cinntaile 5y agoThat's why it's a good thing that the code is open. The usual standard is the same type of code but without public access. This is definitely a step in the right direction.
- epistasis 5y agoAs other comments indicate, how could you be sure the tests are testing the right thing? It's easy enough to slap on a few assert statements (or stopifnot statements, since it's R) about number of rows or something, but that's no replacement for manual inspection of the data. A code-based test can not be invented that will ever substitute for looking at the data in the raw, plotting it, and verification using your full mental capacities. The only way to make sure it's right is the same way you'd do full verification of other code, going through it line by line and making sure it's doing the right thing. Tests can not do that, they can only assert some forms of intent, and they are not good at catching the types of errors that result from ingesting varied data from lots of sources and getting it into modelable form. It's ETL plus a bunch of other stuff going on here.
- checker 5y agoI didn't make the claim that a code-based test can be invented that will substitute for looking at the data in the raw, plotting it, and verification using your full mental capacities. However, "the only way to make sure it's right is to ... go through it line by line" is a bold claim. There are multiple named functions and some unnamed functions in this example that could be verified for programmer mistakes (such as typing 1000 vs 10000) and edge case handling (edge cases that often arise from messy ingested data). But even if they were tested, I'll concede that mistakes can still be made.