4 ms·
IMO, if you're an API company, and you can't handle bursts of traffic from your customers, you should work on improving your backend and stop wasting time messi
by quacker 10y ago
IMO, if you're an API company, and you can't handle bursts of traffic from your customers, you should work on improving your backend and stop wasting time messing around with implementing patterns like this. It's a lose-lose situation for you and your customers.
Improving performance can be significantly more expensive and time-consuming than implementing simple rate limits. A realistic scenario is something along the lines of:
- 99% of time your request rate is N requests/min or lower
- 1% of time your request rate exceeds N requests/min, which could cause service degradation
You can deploy infrastructure to handle the 99% case, slap rate limits in front of the service, and sleep well at night. It's often not worth it to pay for additional infrastructure, spend time optimizing, etc for that 1% case.
As an API user, the way to think about rate limits is as a form of protection from other (misbehaving) users. Everyone is going to be upset when one user's script has a tight-loop firing at 100000 requests/sec and then the API becomes lethargic, throws errors, or goes down all together.
- user5994461 10y agoRequest rate are accounted in requests/second. If you are counting requests/minute, you're seriously averaging your peak load and you're about to epicly fail in production.
- quacker 10y agoYou're being downvoted but there is some truth to this. For example, I inherited a relatively poorly performing service with a per-user limit set to 60 req/min. As the service operator, I have the choice of setting either a rate limit of 60 req/min or 1 req/sec. The former (60/min) leaves you open to per-user spikes of up to 60 req/sec. This is a real risk: 10 users together _could_ produce a spike of 600 req/sec. Still, we went with the per-minute rate limit. Why? Those multi-user bursts are fairly improbable (most of our users were well-behaving) and queueing let us manage these isolated spikes gracefully. The per-minute rate limit is a bit more flexible from a user perspective (reward the good users) and it still combats the problem of sustained heavy load on the service, which is the true danger (stop/limit the "bad" users).
- user5994461 10y agoIt was not talking about the rate limit but the performance of the applications. Performances must never be accounted in request/min. It's only requests/s that matter, because your performances are defined by the peak load you can take (which is many times the minute average). Limits should be on more than a few seconds to support short peaks, yet throttle quickly if they persist.