4 ms·
Queues should have an ideal size based on your latency requirements. Say, you have 100 requests coming in per second. Each consumer of your queue requires 0.3
by edejong 4y ago
Queues should have an ideal size based on your latency requirements.
Say, you have 100 requests coming in per second. Each consumer of your queue requires 0.3 seconds per request. Now, you can optimise for the number of consumers. Would you choose 40 consumers, then the probability of serving the request with an empty queue in 0.3 seconds would be 17% (and > 0.3 seconds would be 84%). However, the total waiting time per request would be 0.1 second.
With 80 consumers, the probability of an empty queue would be higher at 58%. However, most of those consumers are doing nothing, while your waiting time is only reduced by 0.1 second! Only a 30% improvement.
(edit: probability of empty queue in second example is higher, not lower)
- anonymoushn 4y agoCan you state the rest of your assumptions? What is the distribution of the request arrival times?
- edejong 4y agoPoisson. And 1.0 variance in request processing time.
- ThrustVectoring 4y agoI'm not awake/invested enough to make up numbers and do the math here, but shouldn't the cost for provisioning workers and value provided at various latencies matter? Like, each worker you provision costs a certain amount and shifts the latency distribution towards zero by a certain amount, and in principle there should be some function that converts each distribution to a dollar value. Like, suppose you had a specific service level agreement that requires X% of requests to be served within Y seconds over some time interval, and breaking the SLA costs Z dollars. Each provisioning level would generate some distribution of latencies, which can get turned into likelihood of meeting the SLA, and from there you can put a dollar value on it. And crucially, this allows the "ideal" amount of provisioning to vary based off the relative cost of over and under-provisioning; if workers are cheap and breaking the SLA is costly, you would have more workers than if the SLA is relatively unimportant and workers are expensive.
- edejong 4y agoYes, that’s an interesting topic, especially given the prevalence of distributed compute nowadays and the rising awareness of cloud costs. In the end, the distribution is not really poisson, of course. So, you might be interested in low pass filtering to elastically scale your provisioned workers. There is quite some theory about this, including sophisticated machine learned models to predict future load. But I digress.
- deleted 4y ago[deleted]