4 ms·
> We can understand this with the help of Little’s law from queuing theory. It states that the average number of requests in the system (Qd for queue depth) is
by kdwikzncba 3y ago
> We can understand this with the help of Little’s law from queuing theory. It states that the average number of requests in the system (Qd for queue depth) is equal to the average arrival rate of requests (throughput T) multiplied by the average amount of time to serve a request (latency L).
First off, this is obviously false. If you can serve 9req/s and you're getting 10req/s the size of the queue depth is growing at a rate of 1req/s. It's not stationary.
Second, what's the connection between this and gpus? What's the queue? What's the queue depth? What are the requests?
Seems to me that the article focuses more on being smart than actually learning.
- minitoar 3y agoThe scenario is that you’re calculating Qd given a static average latency. Absent that, this formula doesn’t give you a way to compute Qd. What is the average amount of time to service a request in a system where the queue depth is growing without bound?
- ssivark 3y ago> First off, this is obviously false. If you can serve 9req/s and you're getting 10req/s the size of the queue depth is growing at a rate of 1req/s. It's not stationary. I haven’t formally studied any queuing theory, but I think: 1. The rule assumes you have enough processing power to service the average load (otherwise it fails catastrophically like you mentioned) 2. The rule is trying to model the fluctuations in the pending load (which might determine wait time or whatever else).