3 ms·
Not so much "work" per se, but something that I've been thinking about quite a lot recently is how ML practitioners test their systems. For example, before lau
by ml_basics 4y ago
Not so much "work" per se, but something that I've been thinking about quite a lot recently is how ML practitioners test their systems.
For example, before launching a model to prod, I've commonly seen that teams have a suite of datasets/evals, and a model has to be at least as good as the existing model in production before replacing it.
But what about the training pipelines? You can have unit tests for all the individual components, but the analogy of "integration tests" would involve just training a full model, which might be very expensive for large neural networks.
One option would be to have a special tiny dataset and tiny model that can be trained fast on CPU. But this would be an imperfect test of the real system. I've also never seen this in practice, so maybe it wouldn't actually be useful.
Curious if anyone can chime in with their experiences?
- ktrnka 4y agoI once froze a subset of the training data and used it in the test suite for the training pipeline. It was pretty handy as a quick test while refactoring. I don't think I got it integrated into our Jenkins pipeline so it wasn't used by the rest of the team and eventually got outdated as the training data changed. We didn't do the "at least as good" thing though. In prior jobs I'd seen too many legitimate situations in which a metric declines even though there's no regression, such as bug fixes in metrics or occasional updates to the test data. Instead we committed the model evaluation to git and had to review and approve model updates. I wish we'd done more testing for that pipeline, particularly the parts that fetched and preprocessed data. I think we had a couple bugs there over the years, or partial missing data.
- ml_basics 4y agoInteresting, thanks for sharing. > I'd seen too many legitimate situations in which a metric declines even though there's no regression, such as bug fixes in metrics or occasional updates to the test data. I guess you could keep both the old and new versions of the test data or metrics in parallel for a few cycles of model update to get a feel for how the old/new settings differ