3 ms·
Hey Dmitry! Good to meet you -- I'm a big fan of DVC. In our implementation, we've taken the approach of conforming to the standards (typically open-source) se
by yoavz 6y ago
Hey Dmitry!
Good to meet you -- I'm a big fan of DVC. In our implementation, we've taken the approach of conforming to the standards (typically open-source) set by the frameworks for serialization (for example https://www.tensorflow.org/guide/saved_model https://www.tensorflow.org/guide/saved_model). In our API design, it was important to integrate at the framework level (e.g. a tf.keras.models.Model object) for our client libraries. If you're using one of the widely available frameworks that we support, this results in a simple API where the serialization / deserialization is more of an implementation detail. If you're using a custom or rarer ML framework with an unstandardized serialization format, an open-source approach might work better.
Hope that was helpful!
- dmpetrov 6y agoRelying on the frameworks formats seems like a great approach. However, there are a few more open questions related to model governance. Versions of frameworks, input data specification, output data/scores specification. I'd expect all of these pieces to be part of ML model description (this is even more important for model serving, then storing/versioning). It would be great to come up with a common format for all these pieces. So, many levels of ML stack can use it.