35 ms·
Async Python is not faster
- JoeAltmaier 6y agoAsync is about latency. About not stalling the UI thread on I/O etc. Not about speed in the usual sense.
- woofie11 6y agoCooperative multitasking came out slower than preemptive in the nineties, so this is unsurprising in the generic case. I think my question is whether async Python is slower in the case it was designed for -- many, long-running open sockets. Async was traditionally used server-side for things like chat servers, where I might have millions of sockets simultaneously open.
- mpweiher 6y agoYes, the whole hoopla about async and particularly async/await has been a bit puzzling, to say the least. Except for a few very special cases, it is perfectly fine to block on I/O. Operating systems have been heavily optimized to make synchronous I/O fast, and can also spare the threads to do this. Certainly in client applications, where the amount of separate I/O that can be usefully accomplished is limited, far below any limits imposed by kernel threads. Where it might make sense is servers with an insane number of connections, each with fairly low load, i.e. mostly idle, and even in server tasks quality of implementation appears to far outweigh whether the server is synchronous or asynchronous (see attempts to build web servers with Apple's GCD). For lots of connections actually under load, you are going to run out of actual CPU and I/O capacity to serve those threads long before you run out of threads. Which leaves the case of JavaScript being single threaded, which admittedly is a large special case, but no reason for other systems that are not so constrained to follow suit.
- ris 6y ago> Cooperative multitasking came out slower than preemptive in the nineties This wasn't really the reason for the shift away from cooperative multitasking, it was really because cooperative multitasking isn't as robust or well behaved unless you have a lot of control over what tasks you have trying to run together. In theory cooperative multitasking should have better throughput (latency is another story) because each task can yield at a point where its state is much simpler to snapshot rather than having to do things like record exact register values and handle various situations.
- woofie11 6y ago... I never meant to imply that performance was the reason for the switch. We've had a track record of technologies which: 1) Automated things (reliving programmers from thinking about stuff) 2) Were expected to make stuff slower 3) In reality, sped stuff up, at least in the typical case, once algorithms got smart That's true for interpreted/dynamic languages, automated memory management/garbage collection, managed runtimes of different sorts, high-level descriptive languages like SQL, etc. Sometimes, it took a lot of time to figure out how to do this. Interpreters started out an order-of-magnitude or more slower than compilers. It took until we had bytecode+JIT that performance roughly lined up. Then, we started doing profiling / optimization based on data about what the program was actually doing, and potentially aligning compilation to the individual users' hardware, things suddenly got a smidgeon faster than static compilers. There is something really odd to me about the whole async thing with Python. Writing async code in Python is super-manual, and I'm constantly making decisions which ought to be abstracted away for me, and where changing the decisions later is super-expensive. I'd like to write.
- lultimouomo 6y ago> That's true for interpreted/dynamic languages, automated memory management/garbage collection, managed runtimes of different sorts, high-level descriptive languages like SQL, etc. Of the things you mention, I agree on SQL, and "managed runtimes" is generic enough that I cannot really judge. I'm thoughroghly unconvinced about the rest being faster than the alternatives (and that's why you don't see many SQL servers written in interpreted languages with garbage collection).
- woofie11 6y agoWell, I think you missed part of what I said: "at least in the typical case" (which is fair -- it was a bit hidden in there) There's a big difference between normal code and hand-tweaked optimized code. SQL servers are extremely tuned, performant code. Short of hand-written assembly tuned to the metal, little beats hand-optimized C. I was talking about normal apps. If I'm writing a generic database-backed web app, a machine learning system, or a video game. Most of those, when written in C, are finished once they work, or at the very most have some very basic, minimal profiling / optimization. For most code: 1) Expressing that in a high-level system will typically give better performance than if I write it in a low-level system for V0, the stage I first get to working code (before I've profiled or optimized much). At this stage, the automated systems do better than most programmers do, at least without incredible time investments. 2) I'll be able to do algorithmic optimizations much more quickly in a high-level programming language than in C. With a reasonably time-bounded investment in time, my high-level code tends to be faster than my low-level code -- I'll have the big-O level optimizations finished in a fraction of the time, so I can do more of them. 3) My low-level code gets to be faster once I get into a very high level of hand-optimization and analysis. Or in other words, I can design memory management better than the automated stuff, but my get-the-stuff-working level of memory management is no longer better than the automated stuff. I can design data structures and algorithms better than PostgreSQL specific to my use case, but those won't be the first ones I write (and in most cases, they'll be good enough, so I won't bother improving them). Etc.
- aliceryhl 6y agoI would be really interested in seeing some numbers on how long the async code spends before each yield.
- berbc 6y agoIs speed really a good reason for using async? If I remember correctly, asynchronous I/O was introduced to deal with many concurrent clients. Therefore, I would have liked to see how much memory all those workers use, and how many concurrent connections they can handle.
- blondin 6y agosame. maybe author is concerned that many people are jumping the gun on async-await before we all fully understand why we need it at all. and that's true. but that paradigm was introduced (borrowed) to solve a completely different issue. i would love to see how many concurrent connections those sync processes handle.
- calpaterson 6y agoHi - not sure what you mean by this. The sync workers handle one request (to completion) per worker. So 16 workers means 16 concurrent requests. For the async workers it's different - they do more concurrently - but as discussed their throughput is not better (and latency much worse). Maybe what you're getting at is cases where there are a large number of (fairly sleepy) open connections? Eg for push updates and other websockety things. I didn't test that I'm afraid. The state of the art there seems to be using async and I think that's a broadly appropriate usage though that is generally not very performance sensitive code except that you try to do as little as possible in your connection manager code.
- blondin 6y agoyes many open connections is what i meant (suggested by other people as well). by the way, i really liked the writing, it's refreshing. and i agree with you that people aren't using async for the right reasons.
- calpaterson 6y agoThanks :) , really appreciate that. I think all technology goes through a period of wild over-application early on. My country is full of (hand dug) canals for example
- fxtentacle 6y agoNot surprised. The bottleneck in Python tends to be the Global Interpreter Lock. That's also why multithreading only rarely helps and why people attempt multi-process execution instead. But Python is an excellent language for quick prototyping and for controlling other things, like coordinating GPUs who do the actual compute work. So I don't quite get why we need to make Python usable for Webservers, when we already have other languages optimized for that purpose, e.g. Google's Go.
- ris 6y agoThe GIL is not an issue with async python. Async python is single-threaded.
- ghostwriter 6y agoYou can start multiple threads within the same process that don't use an event loop, but run preemptively nonetheless, which will introduce GIL to the thread with an event loop. For instance, you can have a ThreadPoolExecutor running in the same process as the thread with an instantiated event loop.
- ris 6y agoI'm sure you can. And if my aunt had balls, she'd be my uncle.
- ghostwriter 6y agothen she's definitely your uncle and you had better double-check her pronoun, because ThreadPoolExecutor exists for a reason, and it's widely used in pair with run_in_executor() [1], and the pool itself can be shared with other non-executor-related tasks scheduled to run preemptively. [1] https://docs.python.org/3/library/asyncio-eventloop.html#asyncio.loop.run_in_executor https://docs.python.org/3/library/asyncio-eventloop.html#asy...
- dotdi 6y agoEDIT: Read the article again, cleared up why the worker count differs. Anyway, I'm running quite a few small Python services on cheap VPSs, i.e., shit performance, and using async was beneficial for me, with performance being ~30% better. They are bread-and-butter apps that read from Postgres, do some HTTP requests, process the results, and potentially write stuff back to the DB. Same performance gain for other services that have HTTP servers. In my mind, hardware can be used more efficiently with async, since while one routine is waiting for an async result, other routines can run meanwhile.
- pluies 6y agoFrom the article: > Why the worker count varies > The rule I used for deciding on what the optimal number of worker processes was is simple: for each framework I started at a single worker and increased the worker count successively until performance got worse. > The optimal number of workers varies between async and sync frameworks and the reasons are straightforward. Async frameworks, due to their IO concurrency, are able to saturate a single CPU with a single worker process. > The same is not true of sync workers: when they do IO they will block until the IO is finished. Consequently they need to have enough workers to ensure that all CPU cores are always in full use when under load.
- calpaterson 6y ago> Maybe I'm misunderstanding something here, but why is the author benching sync frameworks with 16 workers vs async frameworks with ~5? Hi - that is explained in some detail in the article
- martius 6y agoWhat matters is that the server app uses as much CPU as it can when it needs it. With a sync server, a worker is inactive as long as it is waiting on IO, so you need enough workers to maximize the chance that all workers are busy, else, some clients are waiting even if you've got the CPU to deal with them. With an async server, a single worker handles many clients simultaneously, in theory, a single worker per core is sufficient to eat all the CPU available.
- 6y ago
- jordic 6y agoIt would be pretty nice to see the benchmark with what people is using on the async world (asyncpg + uvloop). Just taking a look on them found it's using aiopog (who is using this?) without uvloop.
- jordic 6y agoAlso on the async use case, it's creating a cursor, when there is no need to use it to just fetch a row from the db (with asyncpg)
- calpaterson 6y agoHi - many of the configurations do use uvloop. For what it's worth, I think people are using aiopg because it works with SQLAlchemy whereas asyncpg does not. I kept the database driver the same because I'm testing sync vs async and not database drivers. I would be interested in testing asyncpg, particularly a performance claim is a big part of that library's documentation but another time.
- jordic 6y agoI don't think people is using aiopg... (at least us). asyncpg also works with sqlalchemy (look at startlette with databases, or our own impl https://github.com/vinissimus/asyncom https://github.com/vinissimus/asyncom ) If I'm not wrong, aiopg it's something not fully async, just because relais on the old driver.
- jordic 6y agothere is also another issue on the benchmarck (not pretty sure how it's handled with aiopg. The pool size. With asyncpg, if you want to handle more req/s you need to increase the pool to at least 15-20 connections, but... this is somehting you need to consider when scaling postgres...
- GamblersFallacy 6y agoHi, a few suggestions. Your benchmarks github repo requirements.txt shows uvloop is not been installed. In addition, the bash script calling uvicorn doesn't have uvloop set for the loop parameter. For example, serve-uvicorn-starlette.sh should be: uvicorn --port 8001 --workers $PWPWORKERS app_starlette:app --loop uvloop The uvicorn docs should point out what a big difference uvloop makes.
- martius 6y agoI'm not sure this is a realistic benchmark. A couple of remarks: 16 workers is not that much considering that modern servers can have a lot of cores available, and I expect that the more workers you need the more likely you'll hit other bottlenecks: * the more workers you need, the more memory you consume (workers are processes, not threads), * I don't know how OS scheduler behave these days, but the general-purpose OS scheduler may consume some CPU time you'd rather give to your app. I understand the point on latency variation though: preemptive multitasking will slice the CPU time "fairly" between workers, while in a cooperative multitasking situation, it would be the job of the programmer to yield after some time.
- calpaterson 6y agoHi - I am confident that 16 workers was the right number for that application deployed on that machine. The machine is described in the article. If you took this app and put it on a machine with 8 cores clearly it would make sense to try 32 workers - but in practice I think few Python apps are so IO bound as this one. Most of the time, just over 2 * cpu count is about the right number. I suspect that scheduler overhead is not a realistic consideration for a Python program. My understanding is that switching executing process takes microseconds at worst, which would be too small to notice from the point of view of a Python programmer. On "it would be the job of the programmer to yield after some time" - I'm always personally suspicious of any technique that rests on programmer diligence. My experience suggests not to require (or even expect!) programmer diligence, even from my own (I assure you, god like) programming abilities. Secondly, yielding more often probably would not help (and in fact I half-suspect part of the problem is the frequent yielding at every async/await keyword!).
- martius 6y agoHi, thanks for your response :) Edit: I've been downvoted so I'll add a precision: Usually, it is believed that async shines against other models once you reach a certain scale (https://en.wikipedia.org/wiki/C10k_problem https://en.wikipedia.org/wiki/C10k_problem). This benchmark shows than async app frameworks are slower than the sync ones when running at a given scale, and since the article doesn't give many details on the incomming traffic, I can only assume that it's low, since it saturates 4 cores. I believe that your conclusion that "Async python is not faster" is an over generalization of your use case. I'm not saying that the configuration in your benchmark is not correct, I am saying that this benchmark may not yield the same results if you try to scale it on bigger hardware. I believe that scheduler overhead can't be ruled out (not for python nor any other program) on a server since we've sometimes observed that the scheduler could be the bottleneck under some circumstances. For instance, some Linux schedulers used to show poor perfs when using nested cgroups with resources quota enabled. Also, I'd like to state my first point again: you need to see how the number of workers will influence the memory usage on your system. Especially with python, if you've got a lot of workers, you can expect some memory fragmentation that can impact the perf of your system.
- kkirsche 6y agoThis reminds me of Rob Pike’s talk from Golang about how concurrency is not parallelism. I think the python community may be hitting this issue where async is meant to model concurrent behavior not always or necessarily facilitate parallel activity
- chooseaname 6y agoI think a good chunk of Python developers expected (expect?) async to be a "get out of GIL free card". It's not.
- compressedgas 6y ago> Function colouring is a big problem in Python Not when you know how to call sync functions from async functions and vice versa. An sync function can call an async function via: loop = asyncio.new_event_loop() result = loop.run_until_complete(asyncio.ensure_future(red(x))) A async function can call a sync function via: loop = asyncio.get_event_loop() result = await loop.run_in_executor(None, blue, x) Where red and blue are defined as: async def red(x): pass def blue(x): pass Note that the documentation is wrong about recommending create_task over ensure_future. That recommendation results in more restrictive code as create_task only accepts a coroutine and not a task. This works for regular functions I don't know how it works for generators.
- ghostwriter 6y agoalternatively, one can use gevent and get a transparent asyncio from a modified runtime - something that a high-level language should've provided out of the box.
- jordic 6y agoHiding awaitables from the language, sounds like against the zen (explicit better than implicit) For example, when someone access a descriptor in Django.. this could end being a query to the db (transparent) but dangerous. With asyncio you explicitly await something to return the execution to the event loop. At least for me sounds like a safer behaviour
- ghostwriter 6y ago> Hiding awaitables from the language, sounds like against the zen (explicit better than implicit) Zen is not respected by explicit asyncio, just try to compose asyncio with iterators [1] [1] https://stackoverflow.com/questions/42448664/async-generator-is-not-an-iterator https://stackoverflow.com/questions/42448664/async-generator... This problem doesn't exist with gevent, and composability is a desired thing in any programming language. Python's asyncio fractioned the community that was previously doing implicit asyncio with sync interfaces, and the current state of API is not an example of composable primitives that follow the Zen of Python: > Beautiful is better than ugly. > Simple is better than complex. > Readability counts. > Special cases aren't special enough to break the rules.
- Thristle 6y agoAny reason why Django wasn't tested? It supports both the sync standard and async stanadard and is AFAIK the most popular web framework (way more then flask)
- Ralfp 6y agoDjango is not async yet. You can run it over ASGI but parts of it (eg ORM) need extra compat layer (sync_to_async wrapper) for async to work.
- Thristle 6y agoStill, should have been tested as a sync framework alongside flask
- VWWHFSfQ 6y agoDjango 3 is only async at the view layer. So if you just want to return a static "Hello World" string then you'll be async. But if you do any I/O then it will block.
- papito 6y agoIt's not? Try this - create 10 Postrges queries and run them in sequence with standard Python. Now yield all those calls asynchronously as an array. What is this even about?
- brianwawok 6y agoNow run them in 10 pythons using multi-proc. Now run then in 10 python threads.
- papito 6y agoExcept that multi-threaded applications are inherently more complex and are extremely challenging to debug.
- brianwawok 6y agoMulti-threading is only hard if you share state. Keep state separate and life is good.
- rkangel 6y agoIt's not about that. It's about 10 different clients all having their connections handled at the same time. There exist mechanisms to do that without async, and this post demonstrates that often they give better results. If one connection needs to do a lot of work (your 10 queries), then that's a different (also important) problem to solve. Async is a nice (from a programmer point of view) way of doing it.
- ChrisMarshallNY 6y agoI use async for UI work, but don't have much of an opinion for servers. I suspect that the best async is that supported by the server OS, and the more efficiently a language/compiler/linker integrates with that, the better. JIT/interpreted languages introduce new dimensions that I have not experienced. I do have some prior art in optimizing libraries, though. In particular, image processing libraries in C++. My opinion is that optimization is sort of a "black art," and async is anything but a "silver bullet." In my experience, "common sense" is often trumped by facts on the ground, and profilers are more important than careful design. I have found that it's actually possible to have worse performance with threads, if you write in a blocking fashion, as you have the same timeline as sync, but with thread management overhead. There are also hardware issues that come into play, like L1/2/3 caches, resource contention, look-ahead/execution pipelines and VM paging. These can have massive impact on performance, and are often only exposed by running the app in-context with a profiler. Sometimes, threading can exacerbate these issues, and wipe out any efficiency gains. In my experience, well-behaved threaded software needs to be written, profiled and tuned, in that order. An experienced engineer can usually take care of the "low-hanging fruit," in design, but I have found that profiling tends to consistently yield surprises. T.A.N.S.T.A.A.F.L.
- unilynx 6y ago> profilers are more important than careful design. > I have found that it's actually possible to have worse performance with threads, if you write in a blocking fashion But isn't excessive blocking/synchronization not something the should already be tackled in your design instead of trying to rework it after the fact ? I would expect profiling to mostly leads to micro-optimisations, eg combining or splitting the time a lock is taken, but when you're still designing you can look at avoiding as much need for synchronization as possible. eg: sharing data copy-on-write (not requiring locks as long as you have a reference) instead of having to lock the data when accessing it. As another commenter says > with asyncio we deploy a thread per worker (loop), and a worker per core. We also move cpu bound functions to a thread pool you can't easily go from eg. thread-per-connection to a worker pool. that should have been caught during design
- 6y ago
- orf 6y agoHis async code creates a pool with only 10 max connections[1] (the default). Whereas his sync pool[2], with a flask app that has 16 workers, has significantly more database connections. I expect upping this number would have a positive effect on asyncio numbers because the only thing[3] this[4] is[5] measuring[6] is how many database connections you have, and is about as far from a realistic workload as you can get. Change your app to make 3 parallel requests to httpbin, collect the responses and insert them into the database. That's an actually realistic asyncio workload rather than a single DB query on a very contested pool. I'd be very interested to see how sync frameworks fare with that. 1. https://github.com/calpaterson/python-web-perf/blob/master/async_db.py#L9 https://github.com/calpaterson/python-web-perf/blob/master/a... 2. https://github.com/calpaterson/python-web-perf/blob/master/sync_db.py#L11 https://github.com/calpaterson/python-web-perf/blob/master/s... 3. https://github.com/calpaterson/python-web-perf/blob/master/app_aio.py#L8 https://github.com/calpaterson/python-web-perf/blob/master/a... 4. https://github.com/calpaterson/python-web-perf/blob/master/app_flask.py#L11 https://github.com/calpaterson/python-web-perf/blob/master/a... 5. https://github.com/calpaterson/python-web-perf/blob/master/app_sanic.py#L11 https://github.com/calpaterson/python-web-perf/blob/master/a... 6. https://github.com/calpaterson/python-web-perf/blob/master/app_starlette.py#L9 https://github.com/calpaterson/python-web-perf/blob/master/a...
- the_mitsuhiko 6y agoHe only has 4 CPUs. I doubt rising the worker count is going to help the async situation. From my experience it’s really hard to make async outperform sync when databases are involved because the async layer adds so much overhead. Only when you are completely io bound with lots of connections does async outperform sync in python.
- orf 6y ago> From my experience it’s really hard to make async outperform sync when databases are involved because the async layer adds so much overhead Highly disagree as the database is just another IO connection to a server, which is asyncio bread and butter. Being able to stream data from longer running queries without buffering and whilst serving other requests (and making other queries) is really quite powerful. But yeah, if you're maxing out your database with sync code then async isn't going to make it magically go faster.
- PythonicAlpha 6y agoVery interesting work and results. It also would be interesting to see the memory footprints of the different solutions.
- willcipriano 6y agoIn my experience, if the workload supports it multiprocess works well in Python. The overhead for a new interpreter is cheap and this sidesteps the GIL issue. If you require communication between your 'threads', ZeroMQ(ØMQ) makes this simple and fast.
- heipei 6y agoNot versed enough in Python and asyncio to replicate or understand his benchmark, however from my simplistic view async (or any concurrent framework really) should almost always be faster with any modern kind of application. Let me give an example: If you run a web-service, chances are that you're gonna make some network call as part of processing a request, be it database, network, search index, etc. With async, these calls free up your application to work on another requests until the call returns. I'd say that many modern web-services are just a stitching-together of external calls (get-user-info-from-db, retrieve-user-items-from-db), so the only real work left for the web application to handle is to wait and parse/encode some JSON to client. If you can't max out your bandwidth in terms of JSON performance with one core you just start as many async processes as you have cores. The big advantage over sync processes then is that your async process can also handle insanely-long-running background calls while still crunching through the other requests. Someone please explain to me if I'm missing something here.
- twic 6y agoThis isn't correct at all. If you have threads, then when one of them does something blocking, it is suspended, which frees up the processor for other threads to do work. It's not like the computer is sitting there idle waiting for IO to happen. If you're writing synchronous code and not using threads, then yes, your analysis is right, but that would be a daft thing to do! The difference between sync and async is mostly whether the state associated with a task (like an incoming HTTP request being serviced) is kept on a dedicated native thread stack (as it is with sync), or in some sort of coroutine structure (as with async). Thread stacks may be somewhat more efficient, but you have to allocate a whole stack upfront for each thread, so if you want to have lots of threads, you need to dedicate a lot of memory to that, and that goes badly. For applications with small numbers of tasks in flight at once, we shouldn't expect a lot of difference between sync and async code. But for tasks with huge numbers of tasks (chat servers are the classic example, but high-traffic webservers with lots of blocking calls in the backend are another), async code should keep chugging on where sync code just falls over. tl;dr async is about the number of tasks you can handle at once, not the speed with which you handle each task.
- 6y ago
- rajandatta 6y agoExcellent article. Well done. Great to see that you examined throughput, latency and other measures. It may not answer all questions that arise in real-life situations and work loads but we need more numerical experiments to really understand how this works.
- deleted 6y ago[deleted]
- 198608_ 6y agoPRO HACKERS HELPING PEOPLE +1302-648-5479 (text) Is your partner keeping secrets of lately and you want to know why? you feel your partner is cheating on you? Do you or someone you know have a police or court case and want the case CLEARED and forgotten by us hacking into FBI or government server and wiping off HISTORY of its existence? Did someone steal your money and you want the person found and your money recovered? Do you feel somebody is spying on you or bugging you and you want the person out of your way or exposed? Did you lost or forget password to your Facebook,Instagram,twitter,Gmail,Yahoomail,Hotmail etc and want them recovered? Do you wish to spy on somebody's computer or phone? Did you loose contact with someone(family member or old friend) and wish to know where they are and how to locate them for you all to reconnect? Did you lose a pet(dog,cat etc)and want them found? You're welcome to our world. We're professional hackers and can invade devices(phones, emails,whasapp,text messages,Facebook,Instagram etc),hack out information you need and forward to you. Then you will stay happy. +13026485479 (texts only) globalhacker1986@gmail.com
- jlg23 6y agoThis debate reminds me on "Node.js Is Bad Ass Rock Star Tech", from 2012: https://www.youtube.com/watch?v=bzkRVzciAZg https://www.youtube.com/watch?v=bzkRVzciAZg
- philote 6y agoIn a post about async web frameworks in Python, I'm really surprised that Tornado was not included.
- yla92 6y agoYeah. I was expecting to see Tornado among the async web frameworks as well. We use it at work for almost all the backend related work and are happy with it.
- kwhitefoot 6y ago> Sadly async is not go-faster-stripes for the Python interpreter. Surely this is back to front. It is go faster stripes because it doesn't make it go faster.
- mikkelam 6y agoTechempower [1] has a really great collection of benchmarks using highly controlled test setups that I like to look at to compare web frameworks. Not affiliated with them, but it's relevant to the post. [1] https://www.techempower.com/benchmarks/#section=data-r19&hw=ph&test=fortune&l=zijzen-1r https://www.techempower.com/benchmarks/#section=data-r19&hw=...
- hombre_fatal 6y agoOne big difference between one thread per request vs single-threaded async code is that synchronization and accessing shared resources is trivial when all of your code is running on a single thread. An entire category of data races like `x += 1` become impossible without you even thinking about it. And that's often worth it for something like a game server where everything is beating on the same data structures. I don't use Python, so I guess it's less of an issue in Python since you're spawning multiple processes rather than multiple threads so you're already having to share data via something out of process like Redis and using its own synchronization guarantees. But for example the naive Go code I tend to read in the wild always has data races here and there since people tend to never go 100% into a channel / mutex abstraction (and mutexes are hard). And that's not a snipe at Go but just a reminder of how easy it is to take things for granted when you've been writing single-threaded async code for a while.
- wwright 6y agoFWIW, Rust gives you the same simplicity (no data races at runtime) with threads as well. (Not necessarily on topic, but if you’re really excited about dodging data races, I figured it would give you something fun to look at!)
- ric2b 6y agoNot in the same way though, it catches the possibility of data races and forces you to rewrite until all the memory accesses are safe. That's more complex to program, you might need to redesign some of your data structures, for example.
- tannhaeuser 6y agoIt's about time someone put this into perspective with figures before more and more people rush to implement business apps in async style (= 80's cooperative multiprocessing). There are exceptions of course; for example Node.js was originally envisioned for eg. game servers where async's purported robustness in the presence of a massive number of open sockets supposedly helps. But I think for the vast majority of workloads going async has a terrible impact to your codebase (either with callback hell or by deprecating most of the host language's flow control primitives like try/catch in favour of hard-to-debug ad-hoc constructs such as Promises). Another price to pay is groking Node.js' streams (streams2/streams3) and domain APIs and unhelpful exception handling story with subtle changes even as late as in v13. As I hear, Python's async APIs aren't uncontroversial either. Now the next thing I'd be interested to get debunked is multithreading vs multiple processes with shared memory (SysV shmem). I'm not very sure, but I'd not been surprised to hear that the predominance of multithreaded runtimes (JVM, most C++ appservers) is purely a cargo-cult effect. As far as I remember, threads were introduced for small and isolated problems in GUI programs, like code completion in IDEs; they were never intended for replacing O/S processes and their isolation guarantees.
- waheoo 6y agoKevlin henney has a lot to say about concurrent processing ithink it was one of thesr talks: https://youtu.be/2yXtZ8x7TXw https://youtu.be/2yXtZ8x7TXw https://youtu.be/ZsHMHukIlJY https://youtu.be/ZsHMHukIlJY Threading is faster, but really only if youre willing to give up your locks and design for it properly.
- rlpb 6y agoI find it interesting that all the talk here is about performance, and nobody has mentioned any benefits of Async Python when performance isn't an issue. I use trio/asyncio to more easily write correct complex concurrent code when performance doesn't matter. See "The Problem with Threads"[1]. For this use case, Async Python probably still isn't faster, but that doesn't matter. Let's not throw out the baby with the bathwater :) [1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2006/EECS-2006-1.pdf https://www2.eecs.berkeley.edu/Pubs/TechRpts/2006/EECS-2006-...
- PaulHoule 6y agoI love asyncio for writing mixed initiative "servers". For instance, I have an asyncio "server" that accepts websocket connections on one side, waits on an AQMP queue, proxies requests and mediates for the HEOS smart speaker API, Phillips Hue, U.S. Weather Service, etc. This is great for react or vue front end applications which get their state updated when things happen in the outside world (e.g. somebody else starts the music player, that gets related) When CPU performance is an issue (say generate a weather video from frames) you want to offload that into another process or thread, but it is an easy programming style if correctness matters.
- O5vYtytb 6y agoThat sounds a lot like home-assistant :)
- PaulHoule 6y agoIt is a little bit, except this one is customizable, maintainable, and not phonish in any way. In particular, there is no "one ring to rule them all" App but rather there are very simple one-task applications (put one button to pair the left/right computer to the soundbar via Optical or Coax) and also some applications that are highly complex (e.g. multiple windows)
- rukittenme 6y agoWhats the point of writing concurrent code if its not faster?
- ronreiter 6y agoNginx on an async server does not make sense.
- VWWHFSfQ 6y agonginx offers many benefits when fronting an application server. for instance, tls termination/offload, request buffering, connection pooling. those are some
- phodge 6y agoHow is this result surprising? The point of coroutines isn't to make your code execute faster, it's to prevent your process sitting idle while it waits for I/O. When you're dealing with external REST APIs that take multiple seconds to respond, then the async version is substantially "faster" because your process can get some other useful work done while it's waiting. Obviously the async framework introduces some overhead, but that bit of overhead is probably a lot less than the 3 billion cpu cycles you'll waste waiting 1000ms for an external service.
- dilandau 6y ago>it's to prevent your process sitting idle while it waits for I/O. ...with the goal of making your application faster.
- arghwhat 6y ago... no. With the goal of allowing concurrency without parallelism. In doing that, you're removing natural parallelism, and end up competing with the kernel scheduler, both in performance and in scheduling decisions.
- parhamn 6y agoThis is a lazy argument. We get it, you know what coroutines are and how the kernel scheduler works (also everyone else in this thread). That doesn't matter though. If you think the average python user is looking for "concurrency without parallelism" with no speed/performance goal in mind, you totally have the wrong demographic. The fact that the language chose to implement asyncio on a single thread (again the end user doesn't care that this is the case, it could have been thread/core abstraction like goroutines), with little gain, which lead to a huge fragmentation of its library ecosystem is bad. Even worse that it was done in 2018. Doesn't matter how smart you are about the internals.
- arghwhat 6y agoHow in the world did you come to the conclusion that I thought Python users wanted that? I simply concluded that it's the only thing it provides. I wasn't saying it was a good thing, which I think was what you might have read it as. Python implements things on a single thread due to language restrictions (or rather, reference implementation restrictions), as the GIL as always disallows parallel interpreter access, so multiple Python threads serve little purpose other than waiting for sync I/O. It's been many years since I followed Python development, but back then all GIL removal work had unfortunately come to a halt...
- Mikhail_K 6y agoHow is it news that Python is slow?
- Spiritus 6y agoFunny how in the TechEmpower Web Framework Benchmarks[1], the async frameworks are basically destroying their sync counterparts. [1] https://www.techempower.com/benchmarks/ https://www.techempower.com/benchmarks/
- hu3 6y agoI'm suprised to see ASP.NET Core taking 1st and 3rd places in the plaintext benchmark. https://www.techempower.com/benchmarks/#section=data-r19&hw=ph&test=plaintext https://www.techempower.com/benchmarks/#section=data-r19&hw=...
- calpaterson 6y agoI think that is largely because they aren't using enough workers. IIRC they use 3 * cpu_count for sync, which just won't be enough. They also misleadingly de-emphasise latency variation. One Python async framework I'd never heard of was top of the pops on throughput there even though the latency numbers suggested it had pretty much fallen apart in the test.
- hedora 6y agoNone of the performant async frameworks in those benchmarks are written in python (unless I missed one). The takeaway of this article is that python’s async io implementations perform poorly. That’s surprising, since async I/O is usually used for performance reasons, and in most other languages, async I/O can be much faster.
- Spiritus 6y agoThe idea is to filter so that it only shows Python. Not to compare with other languages.
- zzzeek 6y agoI am SUPER happy someone else is finally looking at this. It is long past time that the reflexive use of asycnio or systems like gevent/eventlet for no other reason than "hand-wavy SPEED" come to an end. That web applications that literally serve just one user at at time are built in Tornado for "speed". (my example for this is the otherwise excellent SnakeViz: https://jiffyclub.github.io/snakeviz/ https://jiffyclub.github.io/snakeviz/ which IMO should have just used wsgiref). As the blog post apparently cites as well (woo!), I've written about the myth of "async == speed" some years ago here and my conclusions were identical. https://techspot.zzzeek.org/2015/02/15/asynchronous-python-and-databases/ https://techspot.zzzeek.org/2015/02/15/asynchronous-python-a...
- calpaterson 6y agoHi - yes loved your blogpost! Also very tired of the "async magic performance fairy dust" :) It's a difficult myth to dispel and I think the situation in terms of public mindshare is much worse now than it was in 2015. Some very silly claims from the async crowd now have basically widespread credence. I think one of the root causes is that people are sometimes very woolly about how multi-processing works. One of the others is that I think it's easy to make the conceptual mistake of 1 sync workers = 1 async worker and do a comparison that way One of my worries is that right now it feels like everything in Python is being rewritten in asyncio and the balkanisation of the community could well be more problematic than 2 vs 3.
- throwaway894345 6y agoFor me it's worth the effort to deal with async if it means not having to deal with uwsgi or other frontends. But in general I think Python has too many problems (packaging, performance, distribution, etc) that it doesn't make sense IMO to invest in new Python projects.
- 1337shadow 6y agouWSGI is a lot of joy for me, really, I've never been happier with my deployments since I have discovered uWSGI back in 2008 or something, and nowadays it supports plenty of languages so there's just nothing I don't deploy on uWSGI anymore. Python packaging is something that I have fully automated (maintaining over 50 packages here) and that I'm pretty happy with. I fail to see the problem with Python packaging, maybe because I have an aggressive continuous integration practice ? (always integrate upstream changes, contribute to dependencies that I need, and when I'm not doing TDD it's only because I have not yet proof that the code I'm writing is not actually going to be useful) That's not something everybody wants to do (I don't understand their reasoning though). People would rather freeze their dependencies and then cry because upgrading is a lot of work, instead of upgrading at the rhythm of upstream releases. If other packages managers or other languages have packaging features that encourages what I consider to be non-continuous integration then good for them, but that's not how a hacker like me wants to work, being able to "ignore upstream releases" is not a good feature, it made me a sad developer really, "ignoring non-latest releases" have made me a really happy developer. Most performance issues are not imputable to the language. If they are, it's probably not affecting all your features, you can still rewrite the feature that Python is not well performing for into a compiled language. I need most of my code to be easy to manipulate, and very little of it to actually outperform Python. I've recently re-assessed if I should keep going with Python for another 10 years, tried a bunch of languages, frameworks, at the end of the month I still wanted a language that easy to manipulate with basic text tools, that's sufficiently easy so that I can onboard junior collegues on my tools, that provides sufficiently advanced OOP because I find it efficient to structure and reuse code. Python does what it claims, it solves a basic human-computer problem, let's face it: it's here to stay and shine, and its wide ecosystem seems like a solid proof. Wether it makes sense to invest in a project or not should not depend in the language anyway.
- brodouevencode 6y agoGreat article, but don't just abandon async entirely. There are still use cases. For me: I use it to pull data from several external APIs all at once. That data is then married up to produce another data object with some special sauce computation. All of these network calls run in parallel, therefore the network overhead (DNS lookups, SSL handshakes, etc.) all operate at the same time instead of running one after the other if it were in synchronous mode. IIRC the benchmarks for this went from running at 3 minutes+ to just over 20 seconds. So there's still utility, so YMMV.
- Grimm1 6y agoAsync is useful for high IO where you may have a lot of down time between the requests. Are you pulling many requests from different servers with different response times, communicating with a db or pulling out large response bodies. Async is probably going to do better since each one of those synchronously represents potentially large idling periods where other requests could have gotten work done. As to the article the comparisons are good but fails to mention resource constraints, like Gunicorn, forking 16 instances is going to be a lot heavier on memory so for a little more RPS you're probably spending a decent chunk of change more to run your work and I don't think that's worth it considering the Async model in python is pretty easy to grok these days and under this benchmark share a similar performance profile. Now that said If I had to guess these numbers are fine for the average API but if you're doing something like high throughput web crawling or need to serve something on the order of 10's of thousands to hundred thousands RPS async will win out on speed and resource use and ultimately cost. Plus at one point they were like "we could only get an 18% speed up with Vibora" haven't used them my self. But 18% performance increase at really any level of load is fantastic. Hand waving that off tells me the work loads for what is "realistic" don't take in to account real high RPS workloads like you might see at major tech companies.
- meritt 6y ago> forking 16 instances is going to be a lot heavier on memory It really depends on how the application is designed. Fork operates through mmap and copy-on-write. It's extremely lightweight by default. A well-designed fork-based application will already have everything necessary to run a given process into memory, not munge any of the existing shared memory, and only allocate and free memory associated with new events/connections/etc. When programmed that way, individual forks are incredibly light on resources. All the workers are sharing the exact same core application code and logic in memory.
- Grimm1 6y ago"All the workers are sharing the exact same core application code and logic in memory." Oh interesting, are you saying an intelligent forking implementation is able to share static portions of memory with multiple children? I was perhaps under the naive assumption forking was pretty much just a full memory copy of the parent.
- awinter-py 6y agoyeah, having operated a largeish nodejs system that had to be realtime: the latencies were all over the place when it got loaded down MxN (kernel threads to green threads) seems to be the established wisdom for go now, with other langs catching up, but unlike python, go has no GIL so is 'less likely to get stuck' (I say, without a footnote) Wider focus on observability makes me optimistic that we'll do better as an industry at packing software onto hardware, and understanding which workloads benefit from what.
- crimsonalucard1 6y agoNo man. Nodejs will beat the flask benchmark. For this specific test there is no downside to async. What’s going on here is python specific.
- kerkeslager 6y ago1. I don't agree with your conclusion that it's Python specific. You don't have evidence for that--you made that up. And no I'm not interested in whatever benchmark you're going to want to post, because it's not a test of this situation--it cannot possibly be, because when you introduce JS, you're also going to be introducing literally hundreds of other factors which could affect the performance. The assertion you are making is not one you can possibly know. 2. For this specific test, there is a downside to async, as shown by the test. Even if what's going on here were Python-specific (which is still something you made up), downsides to async which only occur in a Python environment are still downsides to async. The title of this post is "Async Python is not faster"--that conclusion is incorrect for many reasons, but none of those reasons include the words "NodeJS", "JS", or anything else that is not in the Python ecosystem. 3. What is going on is probably specific to the tools being used, which is why I said "those downsides certainly don't apply to every project". In fact, they probably don't apply to the idiomatic ways of implementing this in Tornado, for example. But note how I said "probably" because I don't know for sure, and I'm not comfortable with making things up and stating them as facts.
- crimsonalucard1 6y ago>And no I'm not interested in whatever benchmark you're going to want to post, That's rude. Let's put it this way. NodeJS and nginx leveled the playing field. It destroyed the lamp stack and made async the standard way of handling high loads of IO. From that alone it should indicate to you that there is something very wrong with how you're thinking about things. You know the theory of asyncio? Let me restate it for you: If coroutines are basically the SAME thing as routines but with the extra ability to allow tasks to be done in parallel with IO then what does that mean? It means that 5 async workers in theory should be more performant than 5 sync workers FOR highly concurrent IO tasks. The logic is inescapable. So what does it mean, if you run tests and see that 5 async workers are NOT more performant than 5 sync workers ON PYTHON exclusively? The theory of asyncio makes perfect logical sense right? So what is logically the problem here? The problem IS PYTHON. That's a theorem derived logically. No need for evidence or data driven techniques. There's this idea that data drives the world and you need evidence to back everything up. How many data points do you need to prove 1 + 1 = 2? Put that in your calculator 200 times and you got 200 data points. Boom data driven buzzword. That's what you're asking from me btw. A benchmark, a datapoint to prove what is already logical. Then you hilariously decided to dismiss it before i even presented it. Look, I say what I say not from evidence, but from logic. I can derive certain issues about the system from logic. You just follow the logic I gave you above and tell me where it went wrong and why do I need some dumb data point to prove 1+1=2 to you? There is NOTHING made up above. It is pure logic derived from the assumption of what AsyncIO is doing. >But note how I said "probably" because I don't know for sure, and I'm not comfortable with making things up and stating them as facts. But you seem perfectly comfortable in being rude and accusing me of making stuff up. I'm not comfortable in going around the internet and trashing other peoples theories with accusations that they are making shit up. If you disagree say it, I respect that. I don't respect the part where you're saying I'm making stuff up.
- jupp0r 6y agoThe general issue if serving one vs multiple clients per thread has been discussed extensively in the last two decades, see http://www.kegel.com/c10k.html http://www.kegel.com/c10k.html I’m not familiar with python but it seems like there is a glaring performance bug iff using one thread per connection is faster than using async io.
- megaman821 6y agoThreads are quite fast, it shouldn't be that shocking that they are as fast or faster than async. What would be shocking is if threads had less memory usage. So for the c10k problem on a machine with 2 GB of RAM, async will win because threads will exhaust the memory of the machine. Give that same machine 200 GB of RAM and threads may end up being faster.
- jupp0r 6y agoThat’s a very simple model of how memory and cpu interact. If you actually end up using hundreds of Gigabytes of memory, there will be implications on cache hit rates, TLB misses, page table sizes and many other things that make me wary about guessing performance in such a case. There is also the not much discussed issue of having shared resources between all of these threads and the impact of such a threading model on the engineering part if writing such a program. I personally haven’t seen the thread per connection model in a successful large scale server.
- lovasoa 6y agoAsync python is faster when you use it for running parallel tasks. In this benchmark, you are running a single database request per query, so there is no advantage to being asynchronous: a pool of processes will scale just as well (but it will use more memory). The point of async is that it lets you easily make a Postgres query, AND an HTTP query, AND a redis query in parallel.
- antoncohen 6y agoI think this blog post on Python, Gunicorn, and Gevent is relevant: https://rachelbythebay.com/w/2020/03/07/costly/ https://rachelbythebay.com/w/2020/03/07/costly/
- wetmore 6y agoThis blog post is mentioned in TFA
- antoncohen 6y agoYes, sorry I should have been more explicit. The article does mention the Rachel by the Bay blog post, but in a way that makes it sound like a different issue. I think they are more related. I think the Rachel by the Bay blog post does a good job of explaining what event loops are doing under the hood, and how that can lead to bad tail latency for web requests.
- ohyes 6y agoThis is just a fundamental misunderstanding of what concurrency is. I do not see why it requires a benchmark. Concurrency is many things at once. That's it. Async frameworks end up with better concurrency properties because you're not paying the memory and context switching overhead of an entire 'thread' for each thing that you are trying to do at the same time. Instead you are paying the (normally cheaper) overhead of what is essentially a co-routine call. The disadvantage being that you have to manage these context switches yourself, and that they tend to happen more frequently (to maintain the illusion that we are doing many things all at the same time on a single cpu). There is no way that an async framework would ever have better straight line performance than a synchronous one, simply because of all of these extra context switches, and that's fine because that's not what it is for. Imagine I want to have 10,000 requests held open at the same time. Your flask server with 16 workers is going to have a tough time as you don't have enough workers to service that many threads, requests won't get serviced and things will start to time out. Because an async framework multiplexes those workers so that can each individually handle multiple requests at once. Multiplexing in this way costs you something performance wise. If you were to crank up the concurrency beyond 100 at once (the default in the posted scripts), you would start getting different results.
- nojvek 6y agoI don’t understand the benchmark. Different libs have different worker counts. Is the actual throughput test doing any async work? Like calling a db or a REST call before it serves a response? How is this an apples to apples benchmark ?
- mkchoi212 6y ago"the more performance sensitive Python code you can replace the better you will do. This is Python performance tactic with a long history"
- daemonk 6y agoThis seems like a resource allocation issue then? Async python is just starting a bunch of jobs with no regard to how each job claims CPUs. Whereas sync python is using native OS threads which I guess does a much better job of allocating CPUs? For async python, when you make 1000 requests, does it immediately register 1000 jobs across your CPUs via workers for processing? Does that just mean each job takes a tiny piece (1/1000) of the resource pie resulting in slower performance for all jobs? Whereas in sync python you are saying you can only perform X number of jobs at a time where X is the number of allocated workers. So resource allocation is roughly divided into X parts. You also have a db connection pool layer after the server code. Isn't that ultimately your bottleneck? I wonder if your async server is saturating the CPUs making the connection pool slow.
- thayne 6y agoDid the author use an async library to access the database? If not, the benefis of async are diminished by the fact that every request is still synchronously waiting for I/O for the database.
- ericls 6y agoWhat happens if there's a ASGI speaking web server implemented in C/C++/Rust?
- hypewatch 6y agoWhy do you use less than half the workers for the async libraries? uwsggi+flask - 16 workers unicorn+starlette - 5 workers The highest throughout examples in your benchmark all have 16 workers. I also don’t see any hardware data... Does your machine have 16 cores?
- calpaterson 6y agoHi - this is explained in detail in the article.
- hypewatch 6y agoI went straight to the code and didn’t read past the github link - but read through it now. That process described isn’t in the code.. One other thing I noticed is that this uses aiopg for the async db queries instead of asyncpg, which is more widely adopted and IMO much better. I was hoping to re-run these benchmarks myself with asyncpg. Looks like actually running the benchmarks would take a bunch of manual work. In fact I don’t see instructions for running these benchmarks to replicate your results.
- devy 6y agoIMO, Async Python frameworks have been proliferated in the era of AI research projects gaining popularity. Some of these AI researchers who know Python already while working on AI frameworks like pytorch would utilize some of these async python frameworks with cookiecutter [1][2][3][3], which helps many others to quickly spin up a backend api for an AI application, in the names of "fast", both in terms of setup and in terms of responding to requests. [1]: https://github.com/tiangolo/fastapi https://github.com/tiangolo/fastapi [2]: https://github.com/tiangolo/full-stack-fastapi-postgresql https://github.com/tiangolo/full-stack-fastapi-postgresql [3]: https://github.com/tiangolo/uvicorn-gunicorn-fastapi-docker https://github.com/tiangolo/uvicorn-gunicorn-fastapi-docker
- peterthehacker 6y agoI'm trying to figure out how to run these benchmarks on my own machine and experiment with some tweaks to the implementation, but it's unclear how to run these benchmarks from start to finish. I don't see any instructions for running the benchmarks in the github repository. @calpaterson can you provide guidance? I'd like to try an alternative query pattern. The current pattern implemented in the benchmarks is select 1 row in 1 query. I'd like to try an implementation with 2 queries - select count() from table and select * from table limit 10, which is a very common pattern for a REST list view. I would hypothesize that the async apps would perform better in this case, but I'm curious what this benchmark with show.
- calpaterson 6y agoHi - you will need to pip install the requirements into a virtualenv. Then set $PWPWORKERS (eg to 1 to start with) and run serve-gunicorn-flask.sh. That will get you a gunicorn instance up and running. From them on you'll need to set up nginx, pgbouncer and postgres. I used unix sockets between all of these but using TCP/IP is fine. The data generation script is checked in, as is the schema. Before you start you should know that Tudor M (see a PR on the project) experimented with changing the query patterns (to three queries, but not a count(*)). It doesn't change matters and the basic reason for that is that nothing has changed - simply having more blocking or non-blocking IO is irrelevant to throughput - except that the more yields you have the more problematic your response times are going to be under load.
- xvilka 6y agoA good alternative to Python is OCaml. It can be both interpreted with a bytecode, but also compiled natively. With the Multicore OCaml coming[1], along with Domains and algebraic effects, it can be a viable alternative to many cases where Python currently is. Moreover, it offers a strict but flexible typing. [1] https://discuss.ocaml.org/t/multicore-ocaml-may-2020-update/5898 https://discuss.ocaml.org/t/multicore-ocaml-may-2020-update/...
- makz 6y agoLast time I checked async in Python didn’t even work properly, has this changed lately?
- kortex 6y agoHow difficult would it be to write a python runtime, let's call it MPython (meta), that forks into separate Python interpreters, 100% orthogonal, except for channels/shm. Could even use separate entry points, but all the same PID. Multiprocessing (iirc) forks from the first python process (this breaks some things depending on when you fork, eg gRPC with torch.dataloader with multiprocessing will crash). Does that get you anything, or am I misunderstanding how multiprocessing/fork works?
- ryanthedev 6y agoNot a real test. No response time simulation... Learn what async code does for you.
- ectospheno 6y agoDo you have a link handy for a real test?
- birdyrooster 6y agoNo one said it was faster. We said it scaled better. That’s because blocking all execution on IO is bad for time sensitive tasks like web requests. If you want to actually go faster the asyncio interfaces used by aiomultiprocess module get you there by maintaining the event loop across multiple processes. You can save time and memory by sharding your data set and aggregating the return data.
- gotzmann 6y agoThats interesting. In PHP world the situation radically different: modern frameworks based on libevent are really speed up web apps up to 10x. I've thoroughly benchmarked my own framework[1] for REST APIs and now it outperforms many of Go / Node.js platforms on Techempower[2] [1] https://github.com/gotzmann/comet https://github.com/gotzmann/comet [2] https://www.techempower.com/benchmarks/#section=test&runid=e12e0b2d-fc4a-4894-b619-cda198516483 https://www.techempower.com/benchmarks/#section=test&runid=e...
- jordic 6y agoThe benchmark will be better with something like this https://news.ycombinator.com/item?id=12227507 https://news.ycombinator.com/item?id=12227507 On the asyncio side :) Old stories that come again and again
- deleted 6y ago[deleted]
- ahupp 6y agoThis is true as far as it goes, but is not testing the (very common) areas where async shines. Imagine you're loading a profile page on some social networking site. You fetch the user's basic info, and then the information for N photos, and then from each photo the top 2 comments, and for each comment the profile pic of the commentor. You can't just fetch all this in one shot because there's data dependencies. So you start fetching with blocking IO, but that makes your wait time for this request proportional to the number of fetches, which might be large. So instead, you ideally want your wait to be proportional to the depth of your dependency tree. But composing all these fetches that way is hard without the right abstraction. You can cobble it together with callbacks but it gets hairy fast. So (outside of extreme scenarios) it's not really about whether async is abstractly faster than sync. It's about how real developers would solve the same problem with/without async. (Source: I worked on product infrastructure in this area for many years at FB)
- reggieband 6y agoI felt baffled by this thread until I read this response. async/await for me has always been about managing this kind of dependency nightmare. I guess if all you have to do is spawn 100 jobs that run individually and report back to some kind of task manager then the performance gains of threads probably beats async/coroutine based approaches on a pure speed benchmark. But when I have significant chains of dependent work then the very idea of using bare threads and callbacks to manage that is annoying. At least in Typescript nowadays, the ability to just mark a function `async` and throw an `await` in front of its invocation drastically lowers the barrier to moving something from blocking to non-blocking. In the same cases if I had to recommend the same change with thread pools and callbacks (and the manual book-keeping around all that) most developers just wouldn't bother.
- sicromoft 6y ago> just mark a function `async` and throw an `await` ... to [move] something from blocking to non-blocking. That's not how it works. `async` and `await` are merely syntactic sugar around callbacks. Everything in javascript is already nonblocking[1], whether or not you use async/await. [1] There are a few rare exceptions in node js (functions suffixed with "Sync"), but in the same vein, they are blocking whether or not you use async/await.
- bit_logic 6y agoOverall, I think the whole reactive programming style such as node and async python are mistakes at this point. They come at too high a cost for code complexity and maintainability. Synchronous style was always superior with only one flaw, using OS threads. But now there are solutions both existing and upcoming such as Go and Java Project Loom that fix that one flaw. I don't see much appeal in the reactive style at this point.
- rubyn00bie 6y agotldr; Increasing throughput doesn't mean faster, it means more efficient use of your resources. Asynchronous and parallelizing workloads only increases throughput not speed. You get more done faster, you don't get each thing produced faster..er. --- Async is only faster if you're not CPU constrained. I don't think anyone is surprised. The following is really simplified; but, hopefully this makes things more clear for folks... Assuming ONE cpu, with ONE thread Synchronous call: [A: Start]---------------->[B: Finish] Asynchronous call: [A: Start]-------->[B: Pause]...(sleep)...[C: Resume]----->[D: Finish] There is no way to make the async call faster than the synchronous call, period. By simply having the operation pause/wait/resume (context switch) it has introduced overhead that is not present in the synchronous operation. So WTF async? Async is only useful when the context switching overhead is less than the time the I/O operation takes. That's it... So when you have I/O bound tasks, that take more time than it does to switch contexts (and carefully manage how many context you have), you can have increased _throughput_.
- lend000 6y agoI use regular threads in Python3 even for I/O and network requests, just because they are what I am familiar with and I haven't yet found a compelling reason to port projects over that are working fine with the standard thread model (granted, I don't produce that many threads at a time). I haven't noticed any performance hits with the GIL or my shared resource locks, but I'm pretty careful about my concurrent programming. Perhaps someone here can explain what the asyncio paradigm does for you beyond "ease of use" when it doesn't get you past the single processor / GIL issue. In what environments are the "os threads" created by the Python engine actually that expensive? I suppose if you are just starting out it may be easier to grok, but then it won't transfer as well to other programming environments besides perhaps NodeJS.
- jlokier 6y agoI'm not surprised by rhe reuslt. Another commenter said: > async I/O is faster because it avoids context switches and amortizes kernel crossings I think this is widely believed, but it's not particularly true for async I/O (of the coroutine kind meant by async/await in Python, NodeJS and other languages, rather than POSIX AIO). With non-blocking-based async I/O, there are often more system calls for the same amount of I/O, compared with threaded I/O, and rarely fewer calls. It depends on the pattern of I/O how much more. Consider: with async I/O, non-blocking read() on a socket will return -EAGAIN sometimes, then you need a second read() to get the data later, and a bit more overhead for epoll or similar. Even for files and recent syscalls like preadv2(...RWF_NOWAIT), there are at least two system calls if the file is not already in cache. Whereas, threaded I/O usually does one system call for the same results. So one blocking read() on a socket to get the same data as the example above, one blocking preadv() to get the same file data. Every system call is two user<->kernel transitions (entry, exit). The number of these transitions is one of the things we're talking about reducing with async/await style userspace scheduling. Threaded I/O puts all context switches in kernel space, but these add zero user<->kernel transitions, because all the context switches happen inside an existing I/O system call. Another way of looking at it, is async replaces every kernelspace context switche with a kernel entry/exit transition pair instead, plus a userspace context switch. So the question becomes: Does the speed of userspace context switches plus kernel entry/exit costs for extra I/O system calls compare favourably against kernel context switches which add no extra kernel entry/exit costs. If the kernel scheduler is fast inside the kernel, and kernel entry/exit is slow, this favours threaded I/O. If the kernel scheduler is slow even inside the kernel (which it certainly used to be in Linux!), and kernel entry/exit for I/O system calls is fast, it favours async. This is despite userspace scheduling and context switching usually being extremely fast if done sensibly. Everything above applies to async I/O versus threaded I/O and counting user<->kernel transitions, assuming them to be a significant cost factor. The argument doesn't apply to async that is not being used for I/O. Non-I/O async/await is fairly common in some applictions, so that tilts the balance to userspace scheduling, but nothing precludes using a mix of scheduling methods. In fact doing blocking I/O in threads, "off to the side" of an async userspace scheduler is a common pattern. It also doesn't apply when I/O is done without system calls. For example memory-mapped I/O to a device. Or if the program has threads communicating directly without entering the kernel. io_uring is based on this principle, and so are other mechanisms used for communicating among parallel tasks purely in userspace using shared memory, lock-free structures (urcu etc) and ringbuffers.
- alexhutcheson 6y agoA lot of the debate and discussion here seems to come from the fact that the example program demonstrates concurrency across requests (each concurrent request is being handled by a different worker), but no concurrency within each request: The code to serve each request is essentially one straight line of execution, which pauses while it waits for a DB query to return. A more interesting example would be a request that requires multiple blocking operations (database queries, syscalls, etc.). You could do something like: # Non-concurrent approach def handle_request(request): a = get_row_1() b = get_row_2() c = get_row_3() return render_json(a, b, c) # asyncio approach async def handle_request(request): a, b, c = await asyncio.gather( get_row_1(), get_row_2(), get_row_3()) return render_json(a, b, c) # Naive threading approach def handle_request(request): a_q = queue.SimpleQueue() t1 = threading.Thread(target=get_row_1(a_q)) t1.start() b_q = queue.SimpleQueue() t2 = threading.Thread(target=get_row_2(b_q)) t2.start() c_q = queue.SimpleQueue() t3 = threading.Thread(target=get_row_3(c_q)) t3.start() t1.join() t2.join() t3.join() return render_json(a_q.get(), b_q.get(), c_q.get()) # concurrent.futures with a ThreadPoolExecutor def handle_request(request, thread_pool): a = thread_pool.submit(get_row_1()) b = thread_pool.submit(get_row_2()) c = thread_pool.submit(get_row_3()) return render_json(a.result(), b.result(), c.result()) These examples demonstrate what people find appealing about asyncio, and would also tell you more about how choice of concurrency strategy affects response time for each request.
- knite 6y agoThis a great point, surprised you received no follow-up comments!
- danthemanvsqz 6y agoThe benchmark doesn't reflect how I would use asyncio. Instead of simply hitting a DB I'd like to see adding an API call to the middle of the request.
- nDmitry 6y agoJust finished writing a more fare benchmark a few days ago. It's utilizing all cores, have DB pools of the same capacity for all tested languages, uses asyncpg in the async Python version, etc. https://github.com/nDmitry/web-benchmarks https://github.com/nDmitry/web-benchmarks Long story short - asyncio is twice as fast... (results are at the bottom of the readme).