3 ms·
yo congrats again on the launch! Anders dbt Labs here with a "tough" question for you. Apologies for 1) my response being half-baked,and 2) if i haven't done my
by data_ders 4y ago
yo congrats again on the launch! Anders dbt Labs here with a "tough" question for you. Apologies for 1) my response being half-baked,and 2) if i haven't done my homework about Hamilton's features.
coincidentally, my PR to the dbt viewpoint was closed by the docs team as "closed, won't do" [1]
I really like the convention of data plane (where you describe how the data should be transformed) and the control plane (i.e. the configuration of the DAG, do this before this). In this paradigm, I believe that the control plane should be as simple as possible, and even perhaps limited in what can be done with the goal of pushing the user to take data transformation as tantamount. Maybe this is why I fell in love with dbt in the first place is because it does exactly this.
"spicy" take:
allowing users to write imperative code (e.g. using loops) that dynamically generates DAGs are never a good idea. I say this as someone who personally used to pester framework PMs for this exact feature before. While things like task groups (formerly subDAGs) [2] appear initially to be right answer, I always ended up regretting them. They're a scheduling/orchestration solution to a data transformation problem
Can y'all speak to how Hamilton views the data and control plane, and how it's design philosophy encourages users to use the right tool for the job?
p.s. thanks for humoring my pedantry and merging this! [3]
[1]: https://github.com/dbt-labs/docs.getdbt.com/pull/2390 https://github.com/dbt-labs/docs.getdbt.com/pull/2390
[2]: http://apache-airflow-docs.s3-website.eu-central-1.amazonaws.com/docs/apache-airflow/latest/core-concepts/dags.html#taskgroups-vs-subdags http://apache-airflow-docs.s3-website.eu-central-1.amazonaws...
[3]: https://github.com/DAGWorks-Inc/hamilton/pull/105 https://github.com/DAGWorks-Inc/hamilton/pull/105
- krawczstef 4y agoGreat questions! > "spicy" take: allowing users to write imperative code (e.g. using loops) that dynamically generates DAGs are never a good idea. Can you give some examples of when it was a bad idea? Otherwise to clarify, with Hamilton, there is no dynamism at runtime. When the DAG is generated, it's generated and it's fixed. The operator we have for doing this `@parameterize` requires everything to be known at DAG construction time. It's really just short hand for manually writing out all the functions. So I don't think it's quite the same story - it is more a "power user feature" - but when used, it makes code DRY-er, at the cost of some code readability. > While things like task groups (formerly subDAGs) [2] appear initially to be right answer, I always ended up regretting them. They're a scheduling/orchestration solution to a data transformation problem Yep. Hamilton has a concept of `subdag` too. It's really short hand for "chaining" Hamilton drivers. We take the latter approach (I believe) since it is there to help you more easily reuse parts of your DAG with different parameterizations. Since Hamilton isn't concerned with materialization boundaries we don't have to make a decision here so how it impacts scheduling/orchestration can be punted to a later time :) > Can y'all speak to how Hamilton views the data and control plane, Hamilton is just a library. There is no DB that needs to be run to use Hamilton. The only state required is your code. So at the simplest micro-level, Hamilton sits within a task, e.g. creating features, and replaces the python script you'd run there. So at this level, I'd argue data plane vs control plane doesn't really apply, unless you view code as the control plane and where it runs as the data plane... At the macro-level, e.g. a model pipeline pulling data, transforming it, fitting a model, etc., you can logically describe the dataflow with Hamilton, without breaking it up into computational tasks a priori. I'd say Hamilton here tries to be agnostic and provide the hooks you need to help you coordinate your control plane and facilitate operation on your data plane. Note: this is where we see DAGWorks coming in and helping provide more functionality for. E.g. with Hamilton you don't need to decide whether everything runs in a single task say on airflow, or multiple. It's up to you to make that decision. The beauty of which, is that conceptually, changing what is in a task, is really just boilerplate given all the information you have already encoded into your Hamilton DAG. > and how it's design philosophy encourages users to use the right tool for the job? With Hamilton, we believe python UDFs are the ultimate user interface. By using Hamilton we force you to chunk logic, integrations, into functions. We also provide ways to "decorate" function logic which gives the ability to inject logic around the running of said functions. So we're really quite agnostic to the tool, but want to provide the hooks to be able to easily and cleanly add, adjust, remove them. For example, to switch between Ray and Dask, our philosophy is that ideally you can write code that is agnostic to knowing about the implementation. Then at runtime add those concerns in. As another example, the ability to switch/change say observability vendors, should not force a large refactor on your code base. We have an extensible `@check_output` decorator that should constrain how much you "leak" from the underlying tools. In short: (1) write functions that don't leak implementation details, they should just instead try to limit to just expressing logic; (2) the Hamilton framework should have the hooks required for you to plug in "tool" concerns. Does that make sense? Happy to elaborate more.
- slotrans 4y ago> "spicy" take: allowing users to write imperative code (e.g. using loops) that dynamically generates DAGs are never a good idea. Said with all love: you're definitely wrong and it's not hard to demonstrate. 1) some upstream service exports data as a series of chunked files each day, you can't know in advance how many 2) processing a single file pegs a CPU core for ~an hour 3) you want a task per chunk so you can process them in parallel 4) ergo: you NEED dynamic task generation This is a thing that Airflow famously couldn't do until 2.3 or 2.4. Luigi could always do it. Prefect, Nextflow, and Martian(?) can do it. DBT can't do it, though the concept doesn't exactly apply (the DBT version is more like "I want to generate models at compile time" which is much more debatably sane... I have this problem currently and our solution sucks, but mostly I wish I didn't have the problem!) I see where you're coming from, I think, but my experience has been the opposite. In any major project, I absolutely need my DAGs to have a dynamic shape. Forcing them to have a static shape just pushes the required dynamism outside the tool, likely into some nasty code-generation step.
- elijahbenizzy 4y agoInteresting... So I'm not sure I agree. While Hamilton does support this kind of thing (and we're likely going to build out more support), the assumption above is that the unit of work at the task-level is natural to map to individual items in your example. In my experience, you want to decouple the task-level orchestration mechanism from the nature of the data itself. A queuing system with multiple consuming threads/processes or a distributed system like spark is kind of meant for this type of task, and can do it more efficiently. So, why rely on the orchestrator to handle it when you potentially have more sophisticated tooling at hand? To make it more concrete, say your upstream files change to have 10x smaller chunks, and 10x more files -- does the same orchestration system make sense? Are you going to start polluting the set of tasks? If you did want to rely on the orchestrator for parallelism, an alternative strategy could be to chunk and assign the files to each task -- E.G. 3 files per task, and use a stable assignment method to round-robin them (basically the same as the queuing system). Might not get everything done at quite the parallel pace, but the DAG would be stable. Your case is particularly tricky as they take so long to process, but it seems to me that this might be better suited for a listening/streaming tool to react when new files are added, then they can add the semi-processed data to a single location, and your daily orchestration task could read from that and set a high-watermark (basically a streaming architecture). Anyway, not 100% sure how I feel about this but wanted to present a counter-point...