5 ms·
> but if we scale to twice the processing power, then start accepting twice the request rate – we will actually be serving each request in half the time we orig
by mgerdts 2y ago
> but if we scale to twice the processing power, then start accepting twice the request rate – we will actually be serving each request in half the time we originally did.
Not necessarily. If processing power is increased by doubling the clock of a processor or using a hard disk that spins and seeks twice as fast, this may be the case.
But we all know that when you have a single threaded work item, adding a second core does not cause the single threaded work to complete in half the time. If the arrival rate is substantially lower than 1/S, the second core will be of negligible value, and maybe of negative value due to synchronization overhead. This overhead is unlikely to be seen when doubling from 1 to 2, but is more likely at high levels of scaling.
If the processor is saturated, service time includes queue time, and service time is dominated by queue time, increasing processing power by doubling the number of processors may make it so that service time can almost be cut in half. How close to half depends on the ratio of queue time to processing time.
- kqr 2y agoThis is a useful addition. Yes, the above reasoning was assuming that we (a) had located a bottleneck, (b) were planning to vertically scale the capacity of that bottleneck, and (c) in doing so won't run into a different bottleneck! It is still useful for many cases of horizontal scaling because from sufficiently far away, c identical servers looks a lot like a single server at c times the capacity of a single one. Many applications I encounter in practise does not require one to be very far away to make that simplifying assumption. > How close to half depends on the ratio of queue time to processing time. I don't think this is generally true, but I think I see what you're going for: you're using queue length as a proxy for how far away we have to be to pretend horizontal scaling is a useful proxy for vertical scaling, right?