3 ms·
The paper reiterates -- multiple times -- that optimizing for latency is obviously the right choice; the other major signal is instantaneous requests-in-flight
by aseipp 2y ago
The paper reiterates -- multiple times -- that optimizing for latency is obviously the right choice; the other major signal is instantaneous requests-in-flight (RIF). They explicitly mention, on the very first page, that their contributions are not that "low latency is good, aim for that", it's that A) they use a particular combination of RIF and latency to select replicas called their "hot-cold lexicographic rule" which they find works really well, and B) their "probes" (requests by the balancer to the backends, to discover what their current state is) are collected asynchronously rather than in the critical request serving path, to help further drive utilization. Most of the ink in the paper in Section 4 explains the design choices they made around probing as a result.
I'm not done with the paper yet, but the basics are in fact written on page one.