2 ms·
You don't have to send every single request twice, just the ones that are haven't returned in time. Wait until some threshold, such as your p95 latency, and sen
by ak_t 2mo ago
You don't have to send every single request twice, just the ones that are haven't returned in time. Wait until some threshold, such as your p95 latency, and send your backup request after that. Return whichever request comes back first, and it should cut your tail latency without doubling your cost, since it only duplicates the small % of requests at the tail.
Google calls this a 'hedged request': https://cacm.acm.org/research/the-tail-at-scale/ https://cacm.acm.org/research/the-tail-at-scale/