3 ms·
The idea that creating a model can be a “user story written in math” seems to me a variation on a common misunderstanding about what model creation, particularl
by bobbruno 4y ago
The idea that creating a model can be a “user story written in math” seems to me a variation on a common misunderstanding about what model creation, particularly the role of coding in it is. Data scientists, statisticians, modellers, don’t go in knowing what the model is and just coding or specifying it. They use code as a exploratory tool to test several hypotheses until they find one that seems to hold - the code, the data, the algorithm, the statistical test, the exploration and the reasoning are all tools in the process of discovering the right model. It simply can’t be specified in advance, it has to be run and tested.
Having som spaghetti code is a natural consequence of this exploratory, iterative process. Applying good SW engineering practices to this exploratory endeavour is just a natural consequence of the process when you don’t know if your code will be of any use before you finish running and checking the test results. Why would you bother modularising and doing test coverage on something that is very likely to be thrown away after one or 2 runs?
I say this from 28 years of experience both as a data engineer and data scientist. I am a good python developer, and I can write production grade code. But I won’t refactor my code into that until I know that is the code that generates the right model. And I certainly can’t specify this particular code before writing some dirty version of it, testing it and confirming that the model it trains passes some statistical tests, at least.
Basically, the code is not the product - the product is the result of applying some transformations on data and running that through some ML or statistical algorithm to generate a model. Transformations and algorithm being unknown to be useful until tested, hence specification being unknown until coded, run and tested.