3 ms·
Its a bit confusing to claim that "The things your current stack can't give you because it doesn't own the DAG" and use DataBricks as your example: DataBricks i
by mollerhoj 5mo ago
Its a bit confusing to claim that "The things your current stack can't give you because it doesn't own the DAG" and use DataBricks as your example: DataBricks includes jobs and pipelines, so it very much owns the DAG, no?
- hugocorreia90 5mo agoFair point. Databricks owns a scheduling DAG (Workflows, DLT). What I meant by "owns the DAG" is the semantic DAG: model-to-model dependencies with column-level types that the compiler builds. Workflows knows task A runs before task B. Rocky knows `dim_customer.email` flows from `raw_users.email_address` through three CTEs in `stg_customers`. Different layer, same word. I'll be more careful with that framing.
- onlyrealcuzzo 5mo ago> I'll be more careful with that framing. I think you should also try to do a better job selling the benefit of this. As a data engineer, I can see why this might be useful, but glancing through your README, the dots were not completely connected
- hugocorreia90 5mo agoMake sense. Reviewing the README is on my TODO list. Thanks for the heads up!