3 ms·
Luigi is nice, because it is really simple to get started and it gradually allows you to do more complex things like (custom) parameter types, custom targets, e
by mtrn 10y ago
Luigi is nice, because it is really simple to get started and it gradually allows you to do more complex things like (custom) parameter types, custom targets, enhanced super classes and dynamic dependencies, event hooks, task history and more.
One thing that I missed a bit was automatic task output naming based on the parameters of a task, so I wrote a thin wrapper for that [1]. This helps, but mostly for smaller deployments.
Airflow and luigi seemed to me like two side of the same thing: fixed graphs vs data flow. One fixates the DAG, the other puts more emphasis on composition.
That said, I am excited about the data processing tools to come - I believe this is an exciting space and choosing or writing the right tool can make a real difference between a messy data landscape and an agile part of business and business development.
[1] https://github.com/miku/gluish https://github.com/miku/gluish
- jaz46 10y agoDefinitely agree that that is one of the great points with Luigi. Airflow's UI of course blows everything out of the water IMO. As for organizing your data, my personal and very biased opinion is that version control semantics similar to Git [0] are a pretty good way to help tame the complexity of ever-changing data sets. We already version code, but with versioned data too, now everything becomes completely reproducible. [0] Our Data Science Bill of Rights: http://www.pachyderm.io/dsbor.html http://www.pachyderm.io/dsbor.html
- mtrn 10y agoThe git angle would be a huge step forward. What I found is that reproducability is not always on people's mind when they designing such systems, whereas I believe it's one of the most important properties. Pachyderm is on my TODO list for a while, so thanks for reminding me, I'll try to implement something real with it soon.