2 ms·
I did a lot of troubleshooting on a system like this a year or two ago, and most of it came down to making sure that global state transitions are atomic, and ma
by kwillets 8y ago
I did a lot of troubleshooting on a system like this a year or two ago, and most of it came down to making sure that global state transitions are atomic, and making communications as robust as possible.
We had the basics of execute-once using a leasing pattern, but we had a number of bugs related to multiple instances of a task existing in different threads (the executor would load the task object and then fork, leaving two instances in possibly stale states, and I also found failure paths that left multiple instances running), and we also saw a number of daily double-executions related to the lease-renewal process freezing, or non-transactional state transition.
We added a lot of state-transition auditing, including a pid/thread ID to find out where updates were coming from.
IIRC I eventually settled on having the executor (queue listener) do every possible check prior to execution (checking resource limits, process count limits, etc.) without loading the task instance itself (just the ID from the queue message). After the fork the child loads the task and does a single transaction that deletes the queue message and creates the execution record (the one-and-only-one run, basically). Every failure up to that point will requeue, but once the run is created, the queue message has to be deleted. We then transition to leasing the execution, and mark it failed if the lease expires.
We also created a centralized service to renew the leases on the execution objects after we found that to be a failure point. Long-running processes just have a lot of problems keeping connections open, etc.