4 ms·
Queues should absolutely not be empty. If they are often empty, you're over provisioning your consumers. There is an ideal size to a queue, based on the balance
by edejong 4y ago
Queues should absolutely not be empty. If they are often empty, you're over provisioning your consumers. There is an ideal size to a queue, based on the balance equation (https://en.wikipedia.org/wiki/Balance_equation https://en.wikipedia.org/wiki/Balance_equation). Also study queueing theory if you want to know more: https://en.wikipedia.org/wiki/Queueing_theory https://en.wikipedia.org/wiki/Queueing_theory.
- clnq 4y agoThey should be empty if you wish to minimize latency. For example, when emergency room queues are full in hospitals, that is a very undesirable scenario. The same can be applied in software - there are parts of systems that should only use queues to handle spike workloads. And the longer these queues are on average, the worse the overall performance of the system is. Just like patient health outcomes drop with longer ER queues. And really, we need more context to say whether a queue should be empty, trend towards being empty, or towards always having tasks. I don't think it's possible to make a general prescription for all queues.
- edejong 4y agoWhat are spike workloads? Is this poisson? What’s the distribution?
- clnq 4y agoDistributions are useful for optimizing. But the decision whether a queue should generally be empty or have elements comes out of system requirements. Then you can employ statistics to plan what resources you will allocate to processing the queue. Though having capacity for Poisson event spikes isn’t nearly the only way to ensure queues trend to zero elements. For example, you can also scale the workload for each queue element depending on the length of the queue. It’s very popular in video game graphics to allocate a given amount of frame time for ray tracing or anti aliasing, and temporally accumulate the results. So we may process a frame in a real time rendering pipeline for more or less time, depending on the state of the pipeline. Say, if the GPU is still drawing the previous frame, we can accumulate some ray traced lighting info on the CPU for the next frame for longer. If the CPU can’t spit out frames fast enough, we might accumulate ray traced lighting kernels/points for less time each frame and scale their size/impact. Statistical distributions of frame timings don’t figure here much as workloads are unpredictable and may spike at any time - you have to adjust/scale rendering quality on the fly to maintain stable frame rates. This example assumes that we’re preparing one frame on the CPU while drawing the previous one on the GPU. But there may also be other threads that handle frames in parallel. Other workload spikes could be predicted by things like Poisson distributions. But once again, the key takeaway is that you can’t generalise this stuff in any practical application. It all depends on system requirements and context.
- remram 4y agoThey should be empty to reach absolute minimal latency (e.g. the only latency is the time it takes to process the task, which you could call 0 latency). Different services have different latency targets. The right statement is not "they should be empty" but "they should be under the target latency", which might be 0 (then "the queues should be empty"), or might be some other target (then empty means overprovisioned resources). Everyone in this thread seem to be talking over each other because they imagine vastly different services (e.g. shipping company vs hospital emergency rooms).
- bcrl 4y agoIf your queues are empty, you're wasting money paying your cloud / hosting provider too much. What matters is utilization. Once utilization hits ~80%, your queue will grow rapidy. Keep utilization below 80% and you've got a decent tradeoff between throughput and latency. It's a good rule of thumb for networks, and it's a good rule of thumb for services.
- xboxnolifes 4y agoIf you're queue should truly be (not heading toward) empty with the goal of minimizing latency, then you don't need a queue at all. You need more processors.
- highspeedbus 4y agoThis only works in very predictable workloads. Once your average input rate becomes slightly larger than consumption rate your system is screwed. Over provision ensures spikes can be processed along with average rate in a general increased load state for some time.
- edejong 4y agoTrue, my example is an oversimplification. In reality you want to account for rapidly and slowly changing demands. This can be solved with elasticity of systems or by overprovisioning. You can get deep, with time series analysis, low pass filtering and various metrics, such as quantile distributions… however, saying zero queue length is desirable is a gross oversimplification.
- colanderman 4y agoOr use load shedding, in which case you want your queue nonempty to take advantage of negative spikes.
- colanderman 4y agoAdditionally, they must not be empty to maximize throughput of consumers which benefit from batching. If there's only ever one item to grab, consumers cannot amortize fetches across a batch, which impacts throughput, without any inherent benefit to latency. (Imagine sending a TCP packet for every individual byte that was enqueued on a socket!) Same applies if you have any sort of load shedding on the consumer side. If the queue is empty, it means you've shed load you didn't need to; when the next brief drop in load comes, you can't take advantage of it. (Fun analogy: this is why Honda hybrids typically don't charge their battery all the way while driving, to allow the chance to do so for free while braking.)
- avinassh 4y agoAny book recommendations on Queueing Theory?