3 ms·
Optimizing tail latency isn't about saving fleet cost. It's about reducing the chance of one tail latency event ruining a page view, when a page view incurs doz
by pradn 2y ago
Optimizing tail latency isn't about saving fleet cost. It's about reducing the chance of one tail latency event ruining a page view, when a page view incurs dozens of backend requests.
The more requests you have, the higher the chance one of them hits a tail. So the overall latency a user sees is largely dependent on a) number of requests b) tail latency of each event.
This method improves the tail latency for ALL supported services, in a generic way. That's multiplicative impact across all services, from a user perspective.
Presumably, the number of requests is harder to reduce if they're all required for the business.
- tedd4u 2y agoThe last sentence of the abstract indicates the goal is to be able to run at higher utilization without sacrificing latency. The outcome of that is you need less hardware for a given load which for Google undoubtedly means massive cost savings and higher profitability. > Prequal has dramatically decreased tail latency, error rates, and resource > use, enabling YouTube and other production systems at Google to run at much > higher utilization.
- jldugger 2y ago> Optimizing tail latency isn't about saving fleet cost. Indirectly, it is. As the quote I replied to suggests, in order to combat tail latency services often run with surplus capacity. This is just a fundamental tradeoff between the two variables mentioned. So, by improving the LB algo, they (and anyone, really) can reduce the surplus needed to meet any specific SLO.
- pradn 2y agoYou're correct - that's a fair point.