5 ms·
Doesn't 700 requests per second for such a trivial service seem kinda slow?
by snicker7 5y ago
Doesn't 700 requests per second for such a trivial service seem kinda slow?
- fyrn- 5y agoYeah, it's so slow that I'm wondering if they were actually measuring TCP/unix socket overhead. I wouldn't expect to see a difference at such a low frequency.
- Tostino 5y agoYeah, seems like there was some other bottleneck. Maybe changing the IPC method makes the small difference we are seeing, but we should be seeing orders of magnitude higher TPS prior to caring about the IPC method.
- miketheman 5y agoIt may, but this is considering the overhead incurred in the local setup - calling the app through a Docker Desktop exposed port. Running locally on macOS produces ~5000 TPS. The `ab` parameters are also not using any sort of concurrency, and perform requests sequentially. The test is not designed to maximize TPS or saturate the resources, rather isolate variables to perform the comparison.
- hamburglar 5y agoBut this test becomes a lot more interesting when the bottleneck is the actual thing under scrutiny (socket performance), rather than Whatever Python Is Doing. You can probably get 10-20x the RPS out of a trivial golang server, which might shift the bottleneck closer to socket perf.
- miketheman 5y agoPossibly! I wasn't trying to get the highest RPS, rather find a baseline and compare per architecture. I also wrote what I knew how write in a short amount of time. ;)
- lmeyerov 5y agoI didn't look too closely, but I'm wondering if this is Python's GIL. So instead of nginx -> multiple ~independent Python processes, each async handler is fighting for the lock, even if running on a different core. So read as 700 queries / core vs 700 queries / server. If so, in a slightly tweaked Python server setup, that'd be 3K-12K/s per server. For narrower benchmarking, keeping sequential, and doing container 1 pinned to core 1 <-> container 2 pinned to core 2 might tell an even clearer picture. I did enjoy the work overall, incl Graviton comparison. Likewise, OS X's painfully slow Docker impl has driven me to Windows w/ WSL2 -> Ubuntu + nvidia-docker, which has been night and day, so not as surprised that those numbers are weird.
- miketheman 5y agoGlad you liked it! In my example, I'm using an unmodified gunicorn runner to load the uvloop worker. So I'm still only using a single worker process. Once I start tweaking the `--workers` count, I get a much higher queries per second. And you're correct - this is a narrow benchmark, not designed to test total TPS or saturate resources.
- lmeyerov 5y agoGreat, suspected something like that :) Interesting wrt workers -- how does TCP vs sockets start looking in a multiple worker + multicore scenario, esp. wrt peak QPS? That's more confusing for schedulers :) Also, FWIW, the GIL thing might even be true in the case of single core <> single vs multiple workers. Docker supports pinning (`cpuset`?), so should be pretty doable.. We actually have a fairly similar production setup, so up to GPUs confusing things, have been curious on how deep to look in upcoming scaling work. As a fun extra wrinkle, we also used to have nginx as a poor man's api gateway: `request -> [ nginx_container -[tcp]-> app_container1 -[tcp]-> nginx_container -[tcp]-> app_container2 -[tcp]-> ... ]`, but had enough quality issues that we removed the internal nginx indirections.
- miketheman 5y agoWow, that's a lot of indirection! I've had the same setup, but only two requests deep. I think that if there's added value of each nginx tier (if each service has static assets, for example), then there may be some benefit, but tracing through all the layers gets harder. Glad to hear you removed the internal layers!