4 ms·
I've been really interested in creating my own workflow engine for personal use and I kind of had the same thought while trying to plan it out. One approach th
by Ataraxy 6y ago
I've been really interested in creating my own workflow engine for personal use and I kind of had the same thought while trying to plan it out.
One approach that comes to mind to solve this sort of end user issue would be to explicitly define all inputs/outputs for the dataflow of a node as "requires" and "provides". In this manner each node would run in parallel the moment their requirements (ie. dependencies) are met. Additionally, since what each piece of data a node needs and provides is explicitly defined you could technically automatically wire nodes together without having to even really needing know the underlying data model itself.
So it just means needing to define clear unique labels for each port. In a UI a user could just drop nodes into a space which would automatically wire up to matching ports. You can then display which needs are not met and even have an interface for choosing matching nodes that fit what might be missing.
In the end all you would really need to know is how to compose the pieces of logic to get the desired outcome.
I'm a novice at this stuff though so take that with a grain of salt.
- MrSaints 6y agoWhat you have described is quite similar to what Lyft's Flyte is trying to accomplish https://flyte.org/ https://flyte.org/ A lot of Tensorflow inspired DAGs approaches the described node processing in the same way.
- yeswecatan 6y agoYup in theory it makes sense. We do define inputs/outputs for each task and workflow but it's somewhat crude; for instance, we just check if the variable name will be available but not if a specific key is in a dictionary. We can definitely improve this though with schemas (probably json schema) and validation. Wiring this together _without any user input_ (which is our goal) though is very, very hard. Let's say we have two parallel branches-- one that orders boxes with holes in them and another that orders shapes that can fit in those holes. When both orders are in the workflow can join and we want to put the shapes into boxes. How do we know which shape can go in each box? Maybe we have a BoxType with a list of possible shapes. This gets very complicated though when you have many attributes and care about different attributes at different stages of the process. Additionally, if the process should update the database at some point to, say, change a flag from `false` to `true` the user would need to know the underlying data model.
- Ataraxy 6y agoSo if it's a long running task that's definitely a more complex problem that would require some sort of queue/pause/resume system presumably but I imagine the same concept could still apply. The node that merges context would simply have a requirement/dependency on the results "provided" by each branch. It would wait to execute until those requirements were met. The user wouldn't need to know about the data model still, the "merge node" would just intrinsicly wait for the results provided form the seperate branches. Fundamentally when each node completes the system itself just needs to check against the list of nodes to see if its dependencies have just been met and keeps track of which nodes have been already executed so that they don't end up getting triggered again when the next check happens. There will always be a need for some sort of workflow level state or context management that governs all of this orchestration that you would want to persist to a database somewhere if this is a long running workflow but this is a systems concern and the user doesn't need to know about it. That was just a long way of me more or less saying that it doesn't matter how many branches there are, all that matters ultimately is that a node waits to execute when its requirements are met.
- yeswecatan 6y agoI guess what I'm trying to say is that the logic in the merge node is where the magic happens. So you have parallel branches, A and B. A orders boxes while B orders shapes. Somewhere in A and B we created Box and Shape instances and those are outputs of A and B, respectively. The merge node, C, waits until A and B are completed and takes their outputs as input. The merge node needs to know how to match up the Box and Shape instances to say shape 2 can fix in box 1.