3 ms·
I use it quite heavily at my day job. The core pieces of kedro are indispensable on large projects (Pipeline, Catalog, DataSets, Runner, hooks). The marketing
by waylonwalker 5y ago
I use it quite heavily at my day job. The core pieces of kedro are indispensable on large projects (Pipeline, Catalog, DataSets, Runner, hooks). The marketing has a lot of focus on the template, It's a great template, but not the key feature that brings me back to it. Really like that I never have to worry about io, the catalog abstracts this away. I don't worry about long runs to get started with my work. It saves all intermediate steps automatically, and the Pipeline (DAG) lets me run just the parts that I need for my task. Before using kedro I would need to scaffold out a way to save all of these intermediate steps, or suffer from long run times to get started on my task. Lastly the DAG is indispensible during environment migrations, it can quickly tell me all of its edges that I need to be concerned with picking up and moving with me.
Unlike similar projects kedro is just a python framework. It lets me build, deploy, and orchestrate however works best for my team.
- barefeg 5y agoCould you point me to the component responsible for storing the intermediate data of the DAG run? I was looking for this but couldn’t find it from quickly scanning the docs
- joelschw 5y agoHere you go - https://kedro.readthedocs.io/en/0.17.6/09_development/03_commands_reference.html#modifying-a-kedro-run https://kedro.readthedocs.io/en/0.17.6/09_development/03_com...
- waylonwalker 5y agoBy adding a catalog entry for each of your datasets with a filepath argument, it will save them to that place. All their datasets use fsspec under the hood so this path can be an s3 bucket or a local path and it all just works.
- kjkjadksj 5y agoIs this different from snakemake?
- joelschw 5y agoSome similar concepts, but Kedro is specialised on the ML engineering workflow and team collaboration