5 ms·
Wasn't there an article about how the async syntax was benchmarked to actually be slower than the traditional way of using threads? What's the current story on
by leafboi 6y ago
Wasn't there an article about how the async syntax was benchmarked to actually be slower than the traditional way of using threads? What's the current story on python async?
reference: http://calpaterson.com/async-python-is-not-faster.html http://calpaterson.com/async-python-is-not-faster.html
- toxik 6y agoNone of the other replies acknowledge this but it seems you are conflating concurrency and asynchronous. An asynchronous program can be sequentially executed. It is a distinct concept.
- leafboi 6y agoThe async await implementation in python and basic python threading is concurrent under IO. No conflation.
- reticents 6y agoWow, thank you for this link. I appreciate it when my assumptions are challenged like this, particularly given the fact that I have a tendency to take benchmark synopses like FastAPI's [1] for granted. I'll have to be more conscious of the ways in which authors hamstring the competition to game their results. [1] https://fastapi.tiangolo.com/benchmarks/ https://fastapi.tiangolo.com/benchmarks/
- syndacks 6y agoWhy is this being downvoted? Seems like a fair counter-point to me.
- ghostwriter 6y agoI didn't downvote it, but apart from the fact that async io is not meant to be faster (it's all about throughput, after all), the benchmark is flawed and it's been discussed in full before https://news.ycombinator.com/item?id=23496994 https://news.ycombinator.com/item?id=23496994
- leafboi 6y agoasyncio is meant to be "faster" for IO heavy tasks and low compute. The benchmark tests requests per second which is indeed directly testing what you expect it to test. It's been discussed before but the outcome of that discussion (in the link you brought up) was divided. Highly highly divided. There was no conclusion and it is not clear whether the benchmark was flawed. The discussion is also littered with people who don't understand why async is fast for only certain types of things and slow for others. It's also littered with assumptions that the test focused on compute rather than IO which is very very evidently not the case.
- ghostwriter 6y ago> asyncio is meant to be "faster" for IO heavy tasks and low compute. the point is that it's not meant to be any faster than a parallel pool of processes that perform the same heavy IO without blocking all requesting clients. asyncio is about packing as many concurrent socket interactions into a single process as possible, hence optimising for throughput by giving up the speed that gets eaten up by task context-switching. Hence the flaw in the becnhmark. The benchmark was run on the same machine where Postgres was operating. The benchmark used different number of processes for sync and async workloads, the connection pool was not setup to prevent blocking when a coroutine tries to acquire a connection from the pool when the pool is exhausted (for benchmark purposes it should not have upper bound limit and to be pre-populated with already established connections).
- leafboi 6y ago>The benchmark used different number of processes for sync and async workloads. Wrong. Workers amounts are the same. See the chart with benchmark results. http://calpaterson.com/async-python-is-not-faster.html http://calpaterson.com/async-python-is-not-faster.html >The benchmark was run on the same machine where Postgres was operating. This wouldn't affect the variance between sync and async results very much because both frameworks were run on the same machine. >the connection pool was not setup to prevent blocking when a coroutine tries to acquire a connection from the pool when the pool is exhausted. Real world connection pools have an upper bound limit. I don't see why setting an upper bound limit to be closer to reality is not a good test. Also you're completely wrong about the connection pool blocking when it is exhausted. See source code: https://github.com/calpaterson/python-web-perf/blob/master/async_db.py https://github.com/calpaterson/python-web-perf/blob/master/a... If all connections are exhausted then the system still yields to compute and incoming requests. >(for benchmark purposes it should not have upper bound limit and to be pre-populated with already established connections). Disagree. The real world sets an upper bound. There's nothing wrong with simulating this in a test.
- pdonis 6y agoThe "slower" is not really the problem--as the article notes, the sync frameworks it tested have most of the heavy lifting being done in native C code, not Python bytecode, whereas the async frameworks are all pure Python. Pure Python is always going to be slower than native C code. I'm actually surprised that the pure Python async frameworks managed to do as well as they did in throughput. But of course this issue can be solved by coding async frameworks in C and exposing the necessary Python API using bindings, the same way the sync frameworks do now. So the comparison of throughput isn't really fair. The real issue, as the article notes, is latency variation. Because async frameworks rely on cooperative multitasking, there is no way for the event loop to preempt a worker that is taking too long in order to maintain reasonable latency for other requests. There is one thing I wonder about with this article, though. The article says each worker is making a database query. How is that being done? If it's being done over a network, that worker should yield back to the event loop while it's waiting for the network I/O to complete. If it's being done via a database on the local machine, and the communication with that database is not being done by something like Unix sockets, but by direct calls into a database library, then that's obviously going to cause latency problems because the worker can't yield during the database call. The obvious way to fix that is to have the local database server exposed via socket instead of direct library calls.
- leafboi 6y ago>whereas the async frameworks are all pure Python. No it's not pure python. It's a combination. The underlying event loop uses libuv, a C library that's makes up the underlying core of nodejs. The marker of "Uvicorn" is an indicator of this as "Uvicorn" uses uvlib. Overall the benchmark is testing a bit of both. The event loop runs on C but it has to execute a bit of python code when handling the request. >If it's being done via a database on the local machine, and the communication with that database is not being done by something like Unix sockets, but by direct calls into a database library, then that's obviously going to cause latency problems because the worker can't yield during the database call. I am almost positive it is being done with some form of non blocking sockets. The only other way to do this without sockets is to write to file and read from file. There is no "direct library calls" as the database server exists as a separate process to the server process. Here's what occurs: 1. Server makes a socket connection to database. 2. Server sends a request to database 3. database receives request, reads from database file. 4. database sends information back to server. Any library call you're thinking of that's called from the library here may be a "client side" library meaning that the library actually makes a socket connection to the sql server.
- theptip 6y agoThis article doesn't evaluate the case that you actually want ASGI for, so I don't think it's very useful. (Or at least, it confirms something that should have already been clear). If you're compute-bound, then Python async (which uses cooperative scheduling similar to green threads) isn't going to help you. You get parallelism, but not concurrency from this progamming model; only one logical thread of execution is running on the CPU at a time (per-process), so this can only slow you down if you are CPU-constrained. The standard usecase of a sync API backed by a local DB with low request latency is typically going to be compute-bound. This is covered in the Django async docs (https://docs.djangoproject.com/en/3.1/topics/async/ https://docs.djangoproject.com/en/3.1/topics/async/) and also in green threading libraries like gevent (http://www.gevent.org/intro.html#cooperative-multitasking http://www.gevent.org/intro.html#cooperative-multitasking). The case where async workers are interesting is for I/O-bound workloads. Say you're building an API gateway, or your monolithic API starts to need to call out to other API services, particularly external ones like Google Maps API. In this case, the worst-case result is that the proxied HTTP request times out; this could block your Django API's work thread for many seconds. In the async / green-threaded model, this case is fine; you have a green thread/async function call per request, and if that gthread is blocked on an upstream I/O operation, the event loop will just start working on a different API call until the OS gives a response on the network socket. Essentially, there's no reason to use Django async if you're doing a traditional monolithic DB-backed application. It's going to give you benefits in usecases where the standard sync model struggles. (Note, there's an argument that you might want green threads even in a normal monolith, to guard against cases like "developer accidentally wrote a chunky DB query that takes 60 seconds to run for some inputs", but most DB engines don't support one-DB-connection-per-HTTP-connection. There was a bunch of discussion on this topic a few years ago, with the SQLAlchemy author arguing that async is not useful for DB connections: https://techspot.zzzeek.org/2015/02/15/asynchronous-python-and-databases/ https://techspot.zzzeek.org/2015/02/15/asynchronous-python-a... although asyncio support was added: https://docs.sqlalchemy.org/en/14/orm/extensions/asyncio.html https://docs.sqlalchemy.org/en/14/orm/extensions/asyncio.htm...)
- leafboi 6y agoThe tests aren't compute bound. They are testing requests per second. It's testing biased towards IO, not compute. Please read the article.
- tomnipotent 6y agoThe built-in event loop has meh performance, would love to see the benchmarks re-run using libuv - that would help close some of the gap.
- leafboi 6y agoThey max out the speed with tests. It does use libuv. Uvicorn is the indicator as it uses libuv underneath. If you heard of Gunicorn, Uvicorn is the version of Gunicorn with libUV, hence the name.
- calpaterson 6y agoHi, I am the author of the above article. Libuv was used - for example by the uvicorn-based versions.
- bob1029 6y agoI think the story with async is always "it depends", unless we are questioning whether the specific implementation is broken. For some web applications, it might actually be faster (in meaningful aggregate volume) to service a complete request on the calling thread rather than deferring to the thread pool periodically throughout execution. I think the break over point between sync and async comes down to how much I/O (database) work is involved in satisfying the request. If each request only hits the database 1-2 times on average incurring a few milliseconds of added latency, making sync all the way down is might be better than with any amount of added context switching. If each request may take 100-1000 milliseconds to complete overall due to various long-running I/O operations, then async is certainly one good approach for maximizing the number of possible concurrent requests. In most of my applications (C#/.NET Core) I default to async/await for backend service methods, because 9/10 times I am going to the database multiple times for something and I cannot always guarantee that it will return quickly under heavy load. For other items, I explicitly go wide on parallelizable CPU-bound tasks. All of these are handled as a blocking call against a Parallel.ForEach(). Never would a CPU-bound task be explicitly wrapped with async/await, but one may be included as part of a larger asynchronous operation. This stuff used to confuse the hell out of me, and then I finally wrapped my head around the 2 essential code abstractions: async/await for I/O, Parallel.For() (et. al.) for CPU-bound tasks which have parallelism opportunities. Never try to Task.Run or async/await your way out of something that is CPU-bound and is blocking the flow of execution. Try to leverage asynchrony responsibly when delays >1ms are possible in large concurrent volumes.