5 ms·
MLflow v0.8.0 Features Improved Experiment UI and Deployment Tools
- m_ke 8y agoI'm looking into switching over to using MLflow or Polyaxon for experiment management and tracking. We currently us a a custom built django app for experiment tracking and run experiments by hand on desktop workstations but we're starting to move some of that over to GCP. For people who have used either of the projects, what are your opinions and are there any hidden issues that you ran into? Ideally we'd like to have a platform that makes it easy to schedule runs on the desktops or GCP depending on requirements and available resources. Seems like kubernetes might be the best option for that and it doesn't look like MLflow supports it out of the box yet.
- Voloskaya 8y agoPolyaxon is really great in term of functionality and UX. It's still pretty early stage, so there are some bugs, but overall I am very impressed by it. We have been using for a few month now with a couples of ML researchers.
- MostlyAmiable 8y agoMy main issues are that if you're using the serving functionality, the containers it builds take a long time to start because the environment/dependencies are loaded at runtime instead of being baked into the image. Also, it doesn't have the ability to use a db or remote file store to save experiment info, so you need to use EBS volumes or something for persistence.
- mateiz 8y agoWhile MLflow doesn't submit jobs to Kubernetes for you, it should be possible to integrate it with your favorite scheduler to do that. MLflow is designed to accept experiment results from wherever you are running your code, so you can just submit an "mlflow run ..." command to Kubernetes and have it report results to your tracking server.
- manojlds 8y agoWe use it for experiment tracking and model repository in our CI/CD flow. More details on our approach - https://stacktoheap.com/blog/2018/11/19/mlflow-model-repository-ci-cd/ https://stacktoheap.com/blog/2018/11/19/mlflow-model-reposit...
- antisocial 8y agoWe are evaluating MLflow. I would like to know if there are any plans for making this an Apache project?
- mlthoughts2018 8y agoAs an ML engineer, I’ve found MLFlow to be really a disastrously bad way to look at the problem. It’s something that managers or executives buy into without understanding it, and my team of engineers (myself included) have hated it. There are many feature specific reasons, but the biggest thing is that reproduction of experiments needs to be synonymous with code review and the identically same version control system you use for other code or projects. This way reproducibility is a genuine constraint on deployment and deployment of an experiment, whether just training a toy model, incorporating new data, or really launching a live experiment, is conditional on reproducibility and code review of the code, settings, runtime configs, etc., that embodies it totally. This is much better solved with containers, so that both runtime details and software details are located in the same branch / change set, and a full runtime artifact like a container can be built from them. Then deployment is just whatever production deployment already is, usually some CI tool that explains where a container (built from a PR of your experiment branch for example) is deployed to run, along with whatever monitoring or probe tracking tools you already use. You can treat experiments just like any other deployable artifact, and monitor their health or progress exactly the same. Once you think of it this way, you realize that tools like ML Flow are categorically the wrong tool for the job, almost by definition, and they exist mostly just to foster vendor lock-in or support reliance on some commercial entity, in this case Databricks.
- m_ke 8y agoHow is it forcing code review on you? I do agree that having things tied to a commit might not be ideal if you're running a lot of experiments in a large shared codebase. I've been tempted to use git to version my model runs but always avoid it because it's usually just extra work.
- mlthoughts2018 8y agoI think you flipped my comment around. I’m saying that the number one defining requirement of model reproducibility tooling is that it does force version control / code review. It should force the concept of “running an experiment” to be just another instance of a deployment. Any part of running an experiment that happens outside of the scope of that, such as with “mlflow run ...” for example, is immediately violating the most basic property of the whole thing (I guess unless “mlflow run ...” is hacked to perform actual production deployments of all types of programs).