4 ms·
It's doable once you're out of pure experimentation and into the development phase at which test driven development can help. Test that this ETL function expec
by ploika 6y ago
It's doable once you're out of pure experimentation and into the development phase at which test driven development can help.
Test that this ETL function expects a DataFrame with a given schema and returns one with a different (but also known) schema, even with all these edge cases in the filters and group-bys.
Test that the "train_classifier" method/function rejects negative penalisation parameters, returns an object of type X (a trained sklearn object say, or dictionary of weights that can be deserialised), fails loudly if you don't have enough samples from category Y etc.
Test that the predict method returns a probability as a float, a predicted class as an int, a DataFrame with metadata and headers, etc etc.
- barefeg 6y agoThese type tests are not really model tests but of the infrastructure surrounding the model. MLOps tools are tested like normal software development
- Yajirobe 6y agoAnd how do you test that your model has sufficient accuracy? How do you make sure that your model does not deteriorate over time?
- jononor 6y agoassert test_performance > threshold ? This would only be a sanity check for big deviations though. For more subtle changes one may need to do a proper statistical test. I believe the jury is still out of how to do that properly for deep ML models - challenge is the lack of independence in CV folds and generally the compute time it takes to evaluate. However, I would probably not do performance check inside a unit-testing framework. Instead treat this as quality indicators like performance benchmarks, code coverage etc. It may be a "gate", that needs to pass to allow a new model into production. To evaluate performance over time, one would preferably want labeled datasets for test gathered at different points in time. Which requires a (reliable) continuous labeling process. One can also gather customer feedback about performance, track those as metrics. These things are probably more in the "monitoring" part of a system, rather than unit-testing time though.
- profunctor 6y agoYou can test on a tester of use some of the examples from the checklist paper. In that paper they might add "I hate you" to some random data and assert that sentiment doesn't improve.
- ploika 6y agoThose are statistics/ML questions, not software development ones, so software development processes like TDD are only tangentially relevant really. But in any case, it's actually fairly easy to test that your trained model has sufficient accuracy: choose a metric, choose a threshold for said metric, and check that the observed metric on a testing set (data that the model was not trained on) is above the desired threshold. Repeat for several different metrics for a better understanding of how well the model performs. This can be put in a set of unit tests. Check residuals, inspect the logical implications of the regression coefficients (or whatever), plot a few curves etc to be more sure again. This can't really be put in unit tests, nor should it be. But again, this is more statistics than software development. Same goes for model deterioration - every so often you check that the metric(s) still beat the minimum threshold on more recent data.