4 ms·
Similarly, there seems to be partial overlap with MLFlow for tracking iterations. I would find a comparison table vs. existing tools useful, to help me conside
by jpau 6y ago
Similarly, there seems to be partial overlap with MLFlow for tracking iterations.
I would find a comparison table vs. existing tools useful, to help me consider Orchest by placing it in my existing workflow.
- ricklamers 6y agoWe want to try to make it easier for people to understand the landscape of tools and our position within it. I personally like something like GitLab's https://about.gitlab.com/devops-tools/ https://about.gitlab.com/devops-tools/. We'll try to put up something similar on our website at some point. I'm not deeply familiar with MLFlow, but from what I have seen/read it is more of a tracking framework that you can integrate into an existing codebase. While Orchest allows you to take your existing codebase and structure it into a pipeline to get a visual and containerized way of interacting with the codebase (allow a mix of notebooks and .py/.sh/.R scripts), running the pipeline, and visually inspecting success/failure of pipeline runs/steps. Another key point of difference is how we are more concerned with managing the flow of data. Since we let you build pipelines we can give you abstractions to separate data flow from the pipeline code. I.e. letting you define generic pipelines that take any data source (in some standard form, like a schema'd database) and produce reports. Because we control the data source in relation to the containerized pipelines we can also make sure the whole thing performs well when it's being executed in parallel (i.e. same version of the pipeline running grid search over paramaterized pipelines). In other words, we also control more of the underlying infrastructure when executing the pipelines.