4 ms·
I am from McKinsey. I do not work on Kedro, but am a data scientist who has used it. I have mixed feelings on Kedro. Pros: * Forces data scientists to produce
by oneoffmk999 5y ago
I am from McKinsey. I do not work on Kedro, but am a data scientist who has used it. I have mixed feelings on Kedro.
Pros:
* Forces data scientists to produce an end product that is not poorly organized Jupyter notebooks.
* Data Catalog is good for well structured systems
* Pipeline visualization stack is great (Kedro viz)
* Config options are pretty good
* Seems stable. Dev team is pretty good on this and avoiding breaking changes.
Cons:
* Data catalog is kind of bad for any non structured setup with flat file data with manual file movement (which is bad to begin with but sometimes that’s life)
* Productivity of making brand new data science code seems to drop when data scientists leave notebooks and
* Most of the time I get brought into a client context because the client doesn’t know anything about data scientist. A lot of data scientists, from both parties, come from academic backgrounds and aren’t great at code. The nice thing about notebooks is that they run. Kedro requires you to create pipeline and node objects to wrap around your code before it runs. It requires some familiarity with Kedro to understand, run, or modify. This makes it seem like a bad idea to dump on a novice client. If the data scientist on their side inheriting it doesn’t really get it, or leaves, there’s unlikely to be enough internal knowledge to maintain it. I try to avoid pushing any new tech stacks on my clients where I can for this reason.
So… I like it but don’t love it for consulting work, which is ironic.
- sails 5y agoNo pros or cons on deploying models. Seems like that is a significant part of the functional value?
- oneoffmk999 5y agoNot my area of expertise. Can’t comment on it either way. But I would tend imagine it’s a strength of Kedro.
- joelschw 5y agoWill point to some experiences open source users have had on this journey https://medium.com/hacking-talent/production-code-for-data-science-and-our-experience-with-kedro-60bb69934d1f https://medium.com/hacking-talent/production-code-for-data-s... https://medium.com/google-cloud/migrate-kedro-pipeline-on-vertex-ai-fa3f2c6f7aad https://medium.com/google-cloud/migrate-kedro-pipeline-on-ve...
- idomi 5y agoWe've built the Ploomber open source tool for that exact reason - true open source! We've been trying to focus on the data scientists, not taking them out of jupyter and definitely making sure they can execute what they want without a dedicated infra/ops person. Check it out! https://github.com/ploomber/ploomber https://github.com/ploomber/ploomber
- waylonwalker 5y agoMy experience with McK came in two phases. Phase one started with Kedro, and phase 2 was a cash grab, move as hot and fast as you can. In my experience phase 2 got off the ground quickly, but after week the notebook was riddled with run these sections, but not these, 4 devs on the team were hanging out having coffee much of the day because they were waiting for time in the notebook to implement their changes. There was no version management so days were lost to, well someone deleted something they shouldn't have now we need to rewrite it. The lack of code review, linting, and formatting tools left them swimming in messy code that they would clean up later, but later never came. The kedro project is still running nearly 3 years after it started, the notebooks are long forgotten. Notebooks are fine for single contributors if that is what they are comfortable with. If that is what they are comfortable with its probably because they have not experienced engagements like you are bringing to them that require them to collaborate as a larger team. If you plan to effectively run projects that last longer than a few weeks with more than one data scientist, I'd really challenge them to lean into kedro. The long term productivity of the project will greatly benefit from it. You will be showing the team a more sustainable way of creating pipelines that will lead to continued success after you leave. This leaves a better name for you than a quick cash grab that gets long forgotten in my opinion.
- ploomber 5y agoI completely agree with the cons you outlined, especially your point about "productivity drops when data scientists leave notebooks." A few years ago, I started working as a data scientist at a big financial firm and reviewed all workflow orchestrator available tools (including Kedro). I didn't like that all of them forced me to re-write my Jupyter code into their frameworks (they're supposed to make me more productive, not less). True, notebooks have their issues but they can be fixed (I don't buy that "Jupyter is only for prototyping argument"). So, long story short, I started a project with a friend that makes us more productive by fixing the problems that notebooks' problems. https://github.com/ploomber/ploomber https://github.com/ploomber/ploomber
- ricklamers 5y agoReally interesting response. While there's absolutely much to love about Kedro (we spoke to the core dev team - they're a fantastic bunch), we created Orchest to counter the cons you mentioned. https://github.com/orchest/orchest https://github.com/orchest/orchest What do you think of Orchest? Given your unique perspective we'd love to hear what you think of it. We've been having success with agencies, especially when the hand-off needs to be something the clients can easily run with.