2 ms·
There are at least two other reasons I prefer small (paginated) requests. First, it’s much simpler to define meaningful latency SLOs. If requests take roughly
by dap 5y ago
There are at least two other reasons I prefer small (paginated) requests.
First, it’s much simpler to define meaningful latency SLOs. If requests take roughly the same time, I can say that requests (or even requests “of this type”) should complete with p99 under 300ms. If they’re very variable, it’s harder to know when the service is degraded. (You can use time-to-first-byte, which is usually more uniform, but you can miss degradation that affects the data stream. Then it’s tempting to look at throughput, but that’s impacted by client speed, which you usually want to ignore here.)
The other is that bulk requests can make it much cheaper to DoS the service. If all requests take about the same amount of work on the server, then forcing the server to do work requires proportional client resources to make the requests. If you can make one small request that causes the server to go off and do a bunch of work while the client does almost no work, that’s a cheap way to DoS a server.