3 ms·
The way I think about workflow engines is as follows. Please comment and correct me. Keen to discuss. Workflow engines like Cadence essentially work by letting
by mleonard 7y ago
The way I think about workflow engines is as follows. Please comment and correct me. Keen to discuss.
Workflow engines like Cadence essentially work by letting you write regular-looking procedural logic for your workflow. This looks and feels very much like writing an async function with async-await in javascript or C#.
The state in your workflow is then implicit in your code instead of explicit. Here's what I mean by that:
Usually you would explicity serialise your state between each incoming event and do an atomic-compare-and-set operation on an external database to store the new state.
For a new incoming event:
(1) fetch the state,
(2) marshal into an object in your programming language (ie a java class or golang struct),
(3) given the current state, process the event and perform any external actions like sending an email. These external events need to be ok with being done with at-least-once semantics. Update the state object ready for the next incoming event.
(4) serialise the state,
(5) store the state in the database (atomically update with a compare-and-set operation).
(6) Repeat on each event. Do everything with at-least-once repeat-on-failure semantics.
In a workflow engine like cadence, what is persisted to the database is the entire history of events instead of a single state object as described above.
In Cadence the code you write looks very much like async-await style code in languages like javascript or C#. The workflow logic is in some sense an async function that pauses at await statements and picks up again where it paused when the next event comes in.
Remember that cadence stores the entire history of events for a workflow. It does this so that it can rerun the workflow from the beginning, this time with the new incoming event on the end of the history.
Notice that you need to be careful about your workflow being deterministic.
Optimisations:
(1) it knows when it is replaying already-seen-events and doesn't redo external events such as sending emails.
(2) it tries to resend events to the same worker node each time. It caches events at worker nodes.
(3) at the macro level everything is highly-available and repeat-on-failure-with-backoff to ensure progress and at-least-once-semantics.
(4) it supports repeating workflows and child workflows
(5) monitoring, tracing, other things you'd expect
(6) etc
Importantly: notice that there is still in some sense a single state object. The state is just implicit: it is deep down in the internal state machine of the language you wrote the workflow function in (java, golang). Instead of serialising state such as 'time-since-last-email' in a state object to a database... you have 'time-since-last-email' as a local variable in the scope of the workflow function. Similarly your programming language is tracking the call stack and current execution position of the function... normally you'd keep track of progress through the workflow in the state object and condition on this state when receiving a new incoming event.
Thinking about state as explicit (state object approach) versus implicit (replay-history approach) helps me when thinking about cadence and similar workflow engines.
.......................................................
Thanks for reading so far. I'd love to hear from users of cadence at uber or elsewhere:
(1) why do you choose to write workflows with implicit state (by replaying history) instead of storing the state explicitly as a serialised state object in the database?
My guess: developer productivity of writing and maintaining the workflows. Having a common approach and single observable system for many different workflows.
(2) how do you reason about long-running workflows where the business logic needs to be updated? Would this not be much easier if the state object (say a serialised protobuf) was stored explicitly in the database?
(3) wouldn't non-determinism be much easier as well if you stored the state explicitly?
- mfateev 7y ago(1) It simplifies the programming model. There is no way to serialize a state of the call stack through a library in most programming languages. You mentioned that Cadence is similar to C# wait/async. The SWF Flow library is. But the Cadence workflow code is fully synchronous, not requiring callbacks unless needed by the business logic. Applying new events to the cached workfow is also more efficient for large states. (2) It depends. Nothing prevents a workflow writer from checkpointing the state explicitly (by calling continue as new). Infinitely running workflows do it periodically. But having the events history is awesome for rollbacks. For example in Cadence it is possible to rollback a bad change and automatically rollback the state of all your workflows to the good state. In database world a change that corrupted the state is much harder to deal with. (3) The experience shows that the determinism requirement while requires some learning is not that hard to deal with. But the superior programming model it allows is liked by users.