4 ms·
I've used AWS Step Functions extensively over the past several years and give me code every day of the week over the Stepfunctions json config. Once you get bey
by avemg 4y ago
I've used AWS Step Functions extensively over the past several years and give me code every day of the week over the Stepfunctions json config. Once you get beyond a few simple steps, it gets very hard to look at the config and understand what's going on with it. Especially true when you haven't look at the config in awhile. The DAG visualizer definitely helps, but as soon as things get beyond the trivial I long for a different tool.
- coolsunglasses 4y agoI was until a week or two ago part of a team that build datasets with extensive dependencies (thus, complicated DAGs) v1 of the system built before I joined was Step Functions and the like. It gets hairy just as you say. v2 I built and designed with the lead data engineer, we called it Coriaria originally. We're hoping/planning to open source it eventually, although it's a little wrapped around our company's internal needs & systems. It chooses neither "config" strictly speaking nor "code" for the DAG, instead the primary representation/state is all in the PostgreSQL database which tracks the dataset dependencies and how each dataset is built. It's a DAG in PostgreSQL as well. To make dataset creation and management easier, I also wrote a custom Terraform provider for Coriaria. This made migrating datasets into the new system dramatically faster. The provider is really nice, supports `terraform import` and all that. Currently we have it setup so that there are separate roles/accounts that can modify an existing dataset, but reading state only requires authentication, not authorization. This enables one team to depend on another team's dataset as an upstream data source for their datasets without granting permission to modify it or create a potentially stale copy of the dataset. Terraform's internal DAG representation of the resource dependencies is leveraged because "parent_datasets" references the upstream datasets directly, including the ones we don't build. We're able to depend on datasets we don't build ourselves because the system has support for Glue catalog backends to track and register partition availability. Currently, it builds most of the datasets using AWS Athena & S3, however this is abstracted over a single step function. There's no DAG of step functions, it's just a convenient wrapper for the Athena query execution. The system also explicitly understands dataset builds and validations as separate steps. The dashboard makes it easy to trace the DAG and see which datasets are blocking a dataset build. We're adding more integrations to it soon so that other ways of kicking off dataset builds and validations are available. If people are interested in this I can begin lobbying for open sourcing the system. My colleague wanted to open source it as well. All else fails, I'll rebuild it from scratch because I don't like the existing solutions for managing datasets. We've been calling it a data-flow orchestration system or ETL orchestration system, not sure what would be most meaningful to people. I think the main caveat to this system is that I'm not sure how much use it'd be for streaming data pipelines, but it could manage the discretization of streaming into validated partitions wherever streamed data is sunk into. Our operating assumptions are that you want validated datasets to drive business decisions, not raw event data streamed in from Kafka. Making sure the right data is located in each daily (or hourly) partition is part of that validation.
- latchkey 4y agoWhy not just model the json as objects in (insert favorite language) and then use that code to generate the json?
- entropicdrifter 4y agoAh yes, a home-made framework to generate configurations for your framework that's supposed to make your life easier. That way you can maintain your code that maintains your configs that make it easier to run your code that you have to maintain!
- pharmakom 4y agoThis can be a much better approach than upgrading the DAG description language to a true programming language. It forces anything complex to happen at build time where it can do less damage. Plus, we can often use the same library to do static analysis on the output
- savin-goyal 4y agoMetaflow provides a similar concept to interface with Step Functions and Argo Workflows in Python - https://docs.metaflow.org/going-to-production-with-metaflow/scheduling-metaflow-flows/scheduling-with-aws-step-functions https://docs.metaflow.org/going-to-production-with-metaflow/...
- latchkey 4y agoActually, yes. It allows for easier unit and integration testing as well. The original complaint is that things were getting hard to read and they wished there was code for this. It seems logical to create a framework for the json configuration files so that they can be easily mocked and tested. As someone who greatly values spending time on automated testing, it seems weird to not think of it this way. Quick google shows that others have done things like this already... [1] https://noise.getoto.net/2021/10/14/using-jsonpath-effectively-in-aws-step-functions/ https://noise.getoto.net/2021/10/14/using-jsonpath-effective... [2] https://aws.amazon.com/about-aws/whats-new/2022/01/aws-step-functions-support-workflows/ https://aws.amazon.com/about-aws/whats-new/2022/01/aws-step-... [3] https://docs.aws.amazon.com/step-functions/latest/dg/sfn-local-test-sm-exec.html https://docs.aws.amazon.com/step-functions/latest/dg/sfn-loc...
- TYPE_FASTER 4y agoAWS offers a service for managed Airflow: https://aws.amazon.com/managed-workflows-for-apache-airflow/ https://aws.amazon.com/managed-workflows-for-apache-airflow/ Makes me wonder if Amazon internally was using Step Functions, ran into issues trying to scale to larger graphs, realized multiple teams were using Airflow, and created the Managed Airflow service.