3 ms·
I would say that you don't need it if you don't feel hindered in your work. But it is solving a real need. Problems arise when you start to have larger teams wh
by cdavid 6y ago
I would say that you don't need it if you don't feel hindered in your work. But it is solving a real need. Problems arise when you start to have larger teams who interact with models, when the models are applied in production at a regular cadence.
Having your model in docker generated from git is a good first step, but does not solve the most important issues when you are at scale: reproducible training, tracking of experiments including data, ML-specific observability for your models in prod, etc. See e.g. https://martinfowler.com/articles/cd4ml.html https://martinfowler.com/articles/cd4ml.html.
More concretely, since the above easily sounds like a buzzword soup:
1. Experiment-wise, you don't want just want to track your ML model definition, but also track the data and everything else used to build it, not just use it. A typical thing I have seen at every company I worked at: we have this model but we don't know how to reproduce it because the lost the data, or there was some magic numbers in training that may be on an internal wiki if you're lucky.
2. For some important use cases, you want to iteratively work on improving the model in production. That often means work on the data side, feature engineering, tracking skew prod vs training, etc. the model is not often changed. In almost every case, a useful model is a model that sees 100s if not more iterations in production. You need 1. to do 2.
3. Some of those tools are useful to enforce invariant or detect data issues, which is again very common when you run models for a long time. See e.g. tfdev.
But to go back to your point: MLOps is a 2nd order kind of thing. The first order is of course that most ML-related projects are useless, poorly conceived, or even lack any kind of quantitative analysis on the business and/or product. Companies should invest there before MLops IMO.