3 ms·
Sorry, but this is naive and bad advice. Rate limiting is a critically missing component in many service APIs, for more reasons than I can even comprehensively
by jonaf 10y ago
Sorry, but this is naive and bad advice. Rate limiting is a critically missing component in many service APIs, for more reasons than I can even comprehensively enumerate off the top of my head.
Scaling an API of any real value is NOT trivial, and struggling to scale an API to meet user demand does NOT necessarily mean that the backend was poorly designed. This is a naive generalization that is hazardous to the industry. Please don't spread it.
Here are some reasons why a lack of rate limiting / user auth is practically negligence. There are more, to be sure. I have experience operating a customer facing API for Bazaarvoice, so I think I know what I'm talking about. (We do thousands of requests per second and power reviews for the likes of Walmart, Best Buy, and 4,000 other retailers and brands worldwide.)
* Multi-tenancy
* * client A over extends and causes client B to be unable to use the API
* * client A needs scale independent of clients B-F
* monitoring
* * suddenly a client is making fewer than usual calls, why?
* * suddenly a client is making more than usual calls, why?
* billing
* * want more requests / second? Upgrade your contract
* * it's easier to measure how much I should charge customers per request or type of request, when I can see the rates of those requests and what it costs me
* security
* * DDOS attack? Start by setting the limit to nothing, or rejecting the requests
* * leaking API auth info is less dangerous, if it happens
I think some other sibling comments mentioned other great reasons. The takeaway is that a valuable API will most likely be difficult, expensive or both difficult and expensive to scale, and rate limiting is extremely important.
- kpil 10y agoRate limiting, and specifically leaky bucket algorithms that spaces out requests evenly, rather than servicing them as they happens to come in was shown to improve the overall performance in a few systems I worked with. Using a leaky bucket algorithm, and a per-customer bucket, I think it's possible to build "fair" systems that also improves the performance. That is, you can run the system with a higher total transactions per seconds, just by queuing "simultaneous" requests a few milliseconds, as they will complete quicker. The reason is probably that it's reducing contention and levels out the resource usage. I thought it would be a feature in almost all web servers, since it's been known "since forever" in the telecom world, but I have not seen it. (Have not looked specifically either, so maybe there are good support for this everywhere and I missed it...)