3 ms·
I'm quite excited for the emerging Python-, Git-, and DAG-driven data engineering workflows/orchestrators that are code-first, GUI-second. But is anyone daunted
by data_ders 6y ago
I'm quite excited for the emerging Python-, Git-, and DAG-driven data engineering workflows/orchestrators that are code-first, GUI-second. But is anyone daunted at the prospect of investing in one of these platforms, then having to pivot to the paradigm that emerges as the winner 3-5 years from now? I've been trying to keep up with all the tools, but am pretty overwhelmed at this point. Obviouly there's differences, but off the top of my head I can think of:
Airflow, Luigi, Pachyderm, DAGster, Google Cloud Compose, Kuberflow, MLFlow, Azure ML Pipelines, Ascend.io, and so so many more.
- pm90 6y agoThere isn’t a way around it. The best thing to do is to try and build your data pipelines in a way that makes it somewhat easier to swap tools if required, but I’m not sure if there’s a general way to do that.
- jdoliner 6y agoThis is unfortunately always going to be a problem with a developing space. Pachyderm was designed with this issue in mind though. While I don't think we've totally eliminated, or can totally eliminate it, we have done some things to mitigate it. Pachyderm pipelines run your code in a docker container that you define which reads data from the local filesystem. It's designed to be simple and similar to the interface you'd use when playing around with data on your laptop. User's generally have a very easy time migrating existing code into Pachyderm, since often code is already reading from the local filesystem and just needs to be pointed at a new path. Migrating to a different system is normally a little more complicated because it requires integrating whatever that new systems data interface is. Although in another system where you just read from local disk the migration should be quite simple.
- gumby 6y ago> But is anyone daunted at the prospect of investing in one of these platforms, then having to pivot to the paradigm that emerges as the winner 3-5 years from now? This will always be an issue if you’re an early adopter. The three mitigating calculations every case are: 1 - somebody* will probably be the “winner”, and could be it. 2 - between now and 3 years from now will you get enough value that it will be worth migrating to the eventual winner rather than doing without and 3 - three years from now you can just start new projects on the winning tech; most or all of the current projects will have withered and/or died anyway.
- slewis 6y agoShameless plug: my company Weights & Biases (https://wandb.com https://wandb.com) has tackled this in a different way. We make tools to keep track of results across your pipelines regardless of how you choose to orchestrate execution. This is one of our big selling points: very simple on-ramp to reproducible tracking, and infrastructure agnosticism. These have led to broad adoption across companies and academics. We also have a pretty cool UI. Pachyderm makes different tradeoffs and we're excited to see their launch. Seriously congrats to you all! This is an invigorating space to say the least ;).
- jdoliner 6y agoWe have at least one customer who's using both wandb and Pachyderm together. I actually don't think they overlap that much, although I may be misunderstanding that wandb does. Pachyderm doesn't do anything to track models explicitly. Our tracking is specifically about data lineage, i.e. what data was fed into this container to create this result. It looks like wandb does do dataset versioning but it's unclear to me if it's storing references to versions of the dataset or if it's actually the source of truth that's storing the data. I think it's the former but I'm not sure. Pachyderm focuses on the later of being the system that stores and versions the large datasets and presents a unified hash for them. Other systems can then record that hash to have immutable versioned datasets they can rely on. We do have a few philosophical differences though. The biggest being that we're opposed requiring data-scientists to including tracking code to trace their results. Every company we've seen use systems like that winds up with different levels of instrumentation on different pipelines and a lot of experiments with no instrumentation. It makes it hard to answer questions like "what are all the places this data is being used?" conclusively, because there's always this "dark matter" of code that hasn't been instrumented that your code doesn't see. We prefer a system that automatically tracks things without asking the user to do anything so that information is always collected and there when you need it. That being said, I'm not sure how you could do the type of finegrained model tracking you guys do automatically. We can track lineage automatically because we're the system that stores and exposes the data, so we know when it's being accessed.
- slewis 6y ago
- mindhash 6y agoYou are so right. The market is super fragmented. And it feels like there is a new workflow tool everyday. Even I had built one similar to Kubeflow. Finally, I think the survival of tools will come down to being able to compete on the price with leading cloud vendors. Just like the choice of programming languages the decision would get down to your team's past experience and cost of ownership.