3 ms·
To me, it's easiest to think of it as "Make, for data". Your ETL pipelines are often a complex graph of dependencies. If some step halfway down the chain fails,
by abd12 8y ago
To me, it's easiest to think of it as "Make, for data". Your ETL pipelines are often a complex graph of dependencies. If some step halfway down the chain fails, you don't want to restart at the very beginning -- you want to restart where it failed and keep moving.
The ETL frameworks (Airflow, Luigi, now Mara) help with this, allowing you to build dependency graphs in code, determine which dependencies are already satisfied, and process those which are not. They'll usually contain helper code for common ETL tasks, such as interacting with a database, writing to/reading from S3, or running shell scripts.
- martin_loetzsch 8y ago(author here). Mara data integration is indeed a glorified version of Make (with cost based scheduling and lots of visualizations)