4 ms·
Worth highlighting that make defines a DAG, and that you can run make tasks in parallel which will automatically bottleneck (as desired) on common dependencies,
by usgroup 2y ago
Worth highlighting that make defines a DAG, and that you can run make tasks in parallel which will automatically bottleneck (as desired) on common dependencies, and fan out otherwise.
The sort of massive C++ build that make can handle are typically much more complicated than your average ETL pipeline. So, there is plenty of room to grow into make.
- fforflo 2y agoYes, that's extremely important. I've had great success, but replacing Airflow, Luigi, and friends with a cron-ed Makefile target refreshing some database tables (usually Materialized views). I've then used this tool to visualize the execution graph. https://github.com/lindenb/makefile2graph https://github.com/lindenb/makefile2graph The result looks like this: https://tselai.com/data/graph.png https://tselai.com/data/graph.png The convenient thing is that each node in the execution graph is in a different environment. Some are shell scripts, a some are Python scripts while others are SQL queries.
- dspillett 2y agoThere are also a few ways to spread make managed processes over multiple nodes, assuming they share storage, which could be useful if your transforms bottleneck at the CPU rather than on storage or network IO.
- daniel_grady 2y agoThe classics never go out of style. Mike Bostock has a nice article about the use of Make for data workflows: https://bost.ocks.org/mike/make/ https://bost.ocks.org/mike/make/.