4 ms·
The key is that the probes a) are fast: they certainly incur the same network cost of a regular request. But more than that, all they do is read two counters,
by pradn 2y ago
The key is that the probes
a) are fast: they certainly incur the same network cost of a regular request. But more than that, all they do is read two counters, so they're super quick for backends to serve.
b) cheap: they don't do nearly as much work as a "real" request, so the cost of enabling this system is not prohibitive. They simply return two numbers. The probes don't compete with "real" requests for resources.
c) give the load balancer useful information: among all the metrics they could have returned from the backend, the ones they chose led to good prediction outcomes.
One could imagine playing with the metrics used, even using ML to select the best ones, and to adapt them dynamically based on workload and time period.