3 ms·
I'm currently investigating a very similar problem (very high throughput webhooks, specified by customers, to unreliable endpoints) and was considering an archi
by mnutt 8y ago
I'm currently investigating a very similar problem (very high throughput webhooks, specified by customers, to unreliable endpoints) and was considering an architecture involving a set of queues partitioned by webhook response time and/or failure rate.
So if your webhook is bucketed into the 0-100ms queue and your responses start to exceed 100ms, you'd be bumped up to the 100-500ms queue which is more likely to have periodic queuing delays, and upwards from there depending on your response time / failure rate. If your API later recovered and started responding faster you'd be moved back up into the faster queue. That way we could offer different (soft) SLAs for different classes of response times, and scale the workers independently per queue.
I'm curious if there are known issues that people have run into with this approach? The main unknown was how many workers we'd be willing to throw at consistently slow endpoints to try to keep the slower queues from backing up too much, and possibly some flapping as endpoints could respond quickly under low throughput but slow down as soon as they move up to the faster queues.
- lpr22 8y agoI really like this approach. You could even automate it where the 0–100ms queue's workers could sever the connection after 100ms and re-queue the message on the next level up, while also incrementing the timeout counter. You could do really interesting things by incrementing and decrementing a timeout counter... e.g., it gets decremented when the transaction completes under the line for the queue it's in. That could help with flap.