3 ms·
Leave the Gun, Take the Cannoli > When a bad Pub/Sub message was pulled into the binary, an unhandled panic would crash the entire app, meaning web, workers an
by vivegi 3y ago
Leave the Gun, Take the Cannoli
> When a bad Pub/Sub message was pulled into the binary, an unhandled panic would crash the entire app, meaning web, workers and crons all died.
Several years ago while working on a Message Queue based application, we had a poison message, one where the consumer process would read the message and then die because the message was malformed and then the message would get back to the queue and due to the FIFO nature of the queue, the retries will also fail.
Ideally, there should be enough checks to ensure that a poison message can never be created. But if edge cases cause this issue, the handler should have some way of knowing that it is retrying the message for the n-th time and then if it fails, move it to a dead-letter-queue and report failure on handling that message.
Unfortunately, this pattern of issues are quite common. So, we fully expect this and design around it.
Then, any message in a dead-letter-queue gives you an example of a crash-inducing message which can then be debugged for root-cause.
- Arbortheus 3y agoI suppose PubSub is at least slightly better than this because it's not a perfect queue. I.E. messages that are not acknowledged will be retried with exponential backoff. That at least gives some time to process some stuff before you encounter the poison message again. Whereas with a FIFO queue you're completely screwed.