11 ms·
If this effort succeeds (and I hope it does) now Python developers will need to contend with the event-loop albatross of asyncio and all of its weird complexity
by ferdowsi 5y ago
If this effort succeeds (and I hope it does) now Python developers will need to contend with the event-loop albatross of asyncio and all of its weird complexity.
In an alternate Python timeline, asyncio was not introduced into the Python standard library, and instead we got a natively supported, robust, easy-to-use concurrency paradigm built around green/virtual threading that accommodates both IO and CPU bound work.
- harpiaharpyja 5y agoIf you are ever considering making use of asyncio for your project, I would strongly recommend taking a look at curio [1] as an alternative. It's like asyncio but far, far easier to use. [1] https://curio.readthedocs.io/en/latest/index.html https://curio.readthedocs.io/en/latest/index.html
- acidbaseextract 5y agoThe video (or blog post) below is one of the best explanations I've seen about what subtle bugs are easy to make with asyncio, why it's easy to make them, and how the trio library addresses them. But yes, consider alternatives before you pick asyncio as your approach! Talk: https://www.youtube.com/watch?v=oLkfnc_UMcE https://www.youtube.com/watch?v=oLkfnc_UMcE Blog post: https://vorpus.org/blog/notes-on-structured-concurrency-or-go-statement-considered-harmful/ https://vorpus.org/blog/notes-on-structured-concurrency-or-g...
- VWWHFSfQ 5y agoHighly recommend curio
- deleted 5y ago[deleted]
- BiteCode_dev 5y agoWhile the design of Curio is quite interesting, it's may not be a good choice, not for technical reasons, but for logistical reasons: the chances it gets a wide adoption are slim to None. And since we are stuck with colored functions in python, the choice of stack matters very much. Now, if you want easier concurrency, and a solution to a lot of concurrency problems that curio solves, while still being compatible with asyncio, use anyio: https://anyio.readthedocs.io/en/stable/ https://anyio.readthedocs.io/en/stable/ It's a layer that works on top of asyncio, so it's compatible with all of it. But it features the nursery concept from Trio, which makes async programming so much simpler and safer.
- heavyset_go 5y agoanyio is also compatible with asyncio and Trio, so you can use it with either library or paradigm.
- sandGorgon 5y agoUvloop/uvicorn - which is the production grade asgi server only works with asyncio. Hypercorn works with trio..but you lose a LOT of performance
- quietbritishjim 5y agoCurio's spiritual successor is Trio [1], which was written by one of the main Curio contributors and is more actively maintained (and, at this point, much more widely used). Like Curio, it's much easier to use than asyncio, although ideas from it are gradually being incorporated back into asyncio e.g. asyncio.run() was inspired by curio.run()/trio.run(). I have used Trio in real projects and I thoroughly recommend it. This blog post [2] by the creator of Trio explains some of the benefits of those libraries in a very readable way. [1] https://trio.readthedocs.io/en/stable/ https://trio.readthedocs.io/en/stable/ [2] https://vorpus.org/blog/some-thoughts-on-asynchronous-api-design-in-a-post-asyncawait-world/ https://vorpus.org/blog/some-thoughts-on-asynchronous-api-de...
- miohtama 5y agoThis is super interesting. How does async approaches work across different Python libraries? Are there any assumptions about usage of asyncio?
- nerdponx 5y agoThe "async/await" syntax is agnostic of the underlying async library, but Asyncio and Trio provide incompatible async "primitives". So yes, you need to write your code for one or the other, or use Anyio, which is a common layer over both. Many more libraries use Asyncio than Trio, so I have come to recommend Anyio (which has a Trio-like API) but with Asyncio as the backend. That gives you the extensive Asyncio ecosystem but with the structured concurrency design of Trio.
- nerdponx 5y agoAs yet another option, I really like Anyio, which is a "frontend" to either Asyncio or Trio, which provides a consistent and tidy API for what is known as "structured concurrency" (I think the name was popularized by the C library Dill). https://anyio.readthedocs.io https://anyio.readthedocs.io
- tbabb 5y agoWhat specifically is the problem with asyncio? I quite like using it, so I'm curious if there's some aspect that makes it unsustainable?
- calpaterson 5y agoThe key disadvantage is largely that it bifurcates the library base. Async libraries and sync libraries co-exist uneasily in the same program. For nearly every popular library there is now a (usually inferior, less robust) async one. The benefits of Linus' Law are reduced.
- Redoubts 5y agoTrifucates, since now there’s stdlib asyncio, and a popular trio async flavor too.
- tbabb 5y agoFair point. Async is a big enough idea that it probably warrants designing the language with it in mind. I guess another way of phrasing it would be that it violates the "there's only one way to do it" maxim, and the "two ways of doing it" circumstance necessarily came about because the idea was discovered long after the core language and libraries were already written.
- fullstop 5y agoI like using it as well, but I've been bit several times by having runtime exceptions completely swallowed.
- Too 5y agoLet me guess. Using create_task() without await and assuming it will run to completion in the background to trigger some other event? This is almost always the case of async tasks disappearing into the void. After shooting myself in the foot with this one time to many I have a hard rule to never use create_task. Instead make sure every single task ends up in some kind of gather() or wait() awaited from the top level. This will ensure any exceptions are propagated. If the number of tasks are dynamically created, while others are still ongoing, add them to a common list and restart your wait() call.
- BiteCode_dev 5y agoasyncio is not a competition to threads, it's complementary. In fact, it's a perfectly viable strat in python to have several processes, each having several threads, each having an event loop. And it will still be so, once this comes out. You will certainly use threads more, and processes less, but replacing 1000000 coroutines by 1000000 system threads is not necessarily the right strategy for your task. See nginx vs apache.
- dralley 5y agoMultiple threads with one asyncio loop per thread would be absolutely pointless in Python, because of the GIL. With that said, sure, threads and asyncio are complimentary in the sense that you can run tasks on threadpool executors and treat them as if they were coroutines on an event loop. But that serves no purpose unless you're trying to do blocking IO without blocking your whole process.
- bonzini 5y agoIn Python it would be pointless, but for example it's how Seastar/ScyllaDB work: each thread is bound to a CPU on the host and has its own reactor (event loop) with coroutines on it. QEMU has a similar design.
- yellowapple 5y agoIt's also (to my knowledge) how Erlang's VMs (e.g. BEAM) work: one thread per CPU core, and a VM on each thread preemptively switching between processes.
- heavyset_go 5y agoI read it as each process having multiple threads and an event loop. If the threads are performing I/O or calling out to compiled code and releasing the GIL, said GIL won't block the event loop.
- BiteCode_dev 5y agoIt would not be pointless at all, because while one thread may lock on CPU, context switching will let another one deal with IO. This can let you smooth out the progress of each part of your program, and can be useful for workload when you don't want anything to block for too long.
- nine_k 5y agoBTW I wonder why async is so painless in ES6 compared to Python. Why the presence of GIL (which JS also has) did not make running async coroutines completely transparent, as it made running generators (which are, well, coroutines already). Why the whole even loop thing is even visible at all.
- laurencerowe 5y agoBecause JavaScript never had threads so I/O in JavaScript has always been non-blocking and the whole ecosystem surrounding it has grown up under that assumption. JavaScript doesn't need a GIL because it doesn't really have threads. WebWorkers are more akin to multiprocessing than threads in Python. Objects cannot be shared directly across WebWorkers so transferring data comes with the expense of serializing/deserializing at the boundary.
- catlifeonmars 5y agoJS now has shared array buffers.
- laurencerowe 5y agoSharedArrayBuffer is just raw memory similar to using mmap from Python multiprocessing. The developer experience is very different to simply sharing objects across threads.
- BiteCode_dev 5y agoI used them both extensively, and here are the main reasons I can think of: - The event loop in JS is invisible and implicit. V8 proved it can be done without paying a cost for it, and in fact most real life python projects are using uvloop because it's faster than asyncio default loop. JS dev don't think of the loop at all, because it's always been there. They don't have to chose a loop, or thinking about its lifecycle or scheduling. The API doesn't show the loop at all. - Asynchronous functions in JS are scheduled automatically. On python, calling a coroutine function does...nothing. You have to either await it, or pass it to something like asyncio.create_task(). The later is not only verbose, it's not intuitive. - Async JS functions can be called from sync functions transparently. It just returns a Promise after all, and you can use good old callbacks. Instantiating a Python coroutine does... nothing as we said. You need to schedule it AND await it. If you don't, it may or may not be executed. Which is why asyncio.gather() and co are to be used in python. Most people don't know that, and even if you know, it's verbose, and you can forget. All that, again, because using the event loop must be explicit. That's one thing TaskGroup from trio will help with in the next Python versions... - the early asyncio API sucked. The new one is ok, asyncio.run() and create_task() with implicit loop is a huge improvement. But you better use 3.7 at least. And you have to think about all the options for awaiting: https://stackoverflow.com/questions/42231161/asyncio-gather-vs-asyncio-wait https://stackoverflow.com/questions/42231161/asyncio-gather-... - asyncio tutorials and docs are not great, people have no idea how to use it. Since it's more complex, it compounds. E.G, if you use await: With node v14.8+: await async_func(params) With python 3.7+: import asyncio async def main(): # no top level await, it must happen in a loop await async_func(params) asyncio.run(main) # explicit loop, but easy one thanks to 3.7 E.G, deep inside functions calls, but no await: With node: ... async_func(params) With python 3.7+: ... # async_func(params) alone would do nothing res = asyncio.create_task(async_func(params)) ... # you MAY get away with not using gather() or wait() # but you also may get "coroutine is never awaited" # RuntimeWarning: coroutine 'async_func' was never awaited asyncio.gather(res) Of course, you could use "run_until_complete()", but then you would be blocking. Which is just not possible in JS, there is one way to do it, and it's always non blocking and easy. Ironic, isn't it? Beside, which Python dev knows all this? I'm guessing most readers of this post will have heard of it for the first time. Python is my favorite language, and I can live with the explicit loop, but explicit scheduling is ridiculous. Just run the damn coroutine, I'm not instantiating it for the beauty of it. If I want a lazy construct, I can always make a factory. Now, thanks to the trio nursery concept, we will get TaskGroup in the next release (also you can already use them with anyio): async with asyncio.TaskGroup() as tg: tg.start_soon(async_func, params) Which, while still verbose, is way better: - no gather or wait. Schedule it, it will run or be cleaned up. - no need to chose an awaiting strat, or learn about a 1000 things. This works for every cases. Wanna use it in a sync call ? Pass the tg reference in it. - lifecycle is cleanly scoped, a real problem with a lot of async code (including in JS, where it doesn't have a clean solution)
- dekhn 5y agoso true. I've been writing thread-callback code for decades (common in network and gui event loops, see QtPy as an example) and when I looked at asyncio my first thought is "this is not better". It's entirely nontrivial to analyze code using asyncio (or yield) compared to callbacks.
- throwaway81523 5y agoYes, I have a big sense of tragedy about Python 3. Python should run on something like (or maybe the actual) Erlang BEAM with lightweight isolated processes. All my threaded Python code is written using that style anyway (threads communicating through synchronized queues) and I've almost never needed traditional shared mutable objects. Maybe completely never, but I'm not sure about a certain program any more. Added: I don't understand the downvotes. If Python 3 was going to make an incompatible departure from Python 2, they might as well have done stuff like the above, that brought real benefits. Instead they had 10+ years of pain over relatively minor changes that arguably weren't all improvements.
- BeetleB 5y agoYou are likely being downvoted because most claims about the pain of a Python 3 transition are inflated/hyperbole. It took less than a day to migrate all my code to Python 3. And by "less than a day" I mean "less than 2 hours". Granted, bigger projects would take longer, but saying stuff like "10+ years of pain" is ridiculous. Probably less than 1% of projects had serious issues with the migration. We just hear of a few popular ones that had some pain and assume that was representative.
- throwaway81523 5y agoThe entire Python community was in pain over Python 3 for 10 years, even if migrating any particular program wasn't much trouble. If you want to contest the notion that there was pain, then fine: most of the community simply ignored Python 3 for 10 years, because there was no reason until quite late in the process to worry about it. I myself never bothered migrating any of my Python 2 stuff. It might not be difficult to do so, but continuing to run it under Python 2 still works fine. If you migrated all of yours in 2 hours, you must not have had much to start with. I do use Python 3 for new stuff most of the time by now, but I keep running into little snags, like the .decode() method not working on strings, or having some (but not all) of the codecs removed from the string module so you have to use the codecs module. There's also the matter of stuff that is supposedly ported but isn't completely. For example, Beautiful Soup works nicely under py3, but it doesn't have a typeshed entry so its import needs a special annotation to stop mypy from complaining about it. The real loss with Python 3 is that it could have been so much better than it is. I remember hearing that Go expected to pick up a lot of migrating C++ users, but it got migrating Python users instead. Here's a pain story about a 2 to 3 migration though: https://dropbox.tech/application/how-we-rolled-out-one-of-the-largest-python-3-migrations-ever https://dropbox.tech/application/how-we-rolled-out-one-of-th...
- KaiserPro 5y ago> easy-to-use concurrency paradigm Well it has queues and threads already. Its just that asyncio for socket handling at least (in the testing that I did) is about 5% faster. (one asyncio socket "server" vs ten threads [with a number of ways to monitor for new connections]) I always assumed that people wanted asyncio because they look at javascript and thought "hey I want GOTOs cosplaying as a fun paradigm"
- BiteCode_dev 5y agoGOTO cosplaying should go away with structured concurrency (via TaskGroup) being adopted in 3.11, as pioneered by Trio. Check out anyio if you want to use them now.
- quietbritishjim 5y agoTaskGroup in asyncio has been promised at least as far back as Python 3.8 [1]. They're still not in the draft "What's new in python 3.11" [2] and searching the web didn't return any official statements. I believe they're planned but don't believe they'll arrive any time soon. If you want to use structured concurrency now, IMO the best bet is to use Trio directly. Reading posts by the author makes it clear that every detail of the library had been extremely scrutinised, not just the API (e.g. see this long post on ctrl-C handling [3], or any number of long technical discussions on the issue tracker), so I think it's a better choice in any case. [1] https://twitter.com/1st1/status/1041855365745455104 https://twitter.com/1st1/status/1041855365745455104 [2] https://docs.python.org/3.11/whatsnew/3.11.html https://docs.python.org/3.11/whatsnew/3.11.html [3] https://vorpus.org/blog/control-c-handling-in-python-and-trio/#other-async-libraries https://vorpus.org/blog/control-c-handling-in-python-and-tri...
- BiteCode_dev 5y agoAnyio is probably a best bet to use TaskGroup IMO. You can use the trio backend if you want, after all. But it's asyncio compatible, which makes it more future proof.
- btown 5y ago> instead we got a natively supported, robust, easy-to-use concurrency paradigm built around green/virtual threading that accommodates both IO and CPU bound work Minus the "natively supported" part, we have this today in http://www.gevent.org/ http://www.gevent.org/ ! It's so, so empowering to be able to access the entire historical body of work of synchronous-I/O Python libraries, and with a single monkey patch cause every I/O operation, no matter how deep in the stack, to yield to your greenlet pool without code changes. We fire up one process per core (gevent doesn't have good support for multiprocessing, but if you're relying on that, you're stuck on one machine anyways), spend perhaps 1 person-day a quarter dealing with its quirks, and in turn we never need to worry about the latencies of external services; our web servers and batch workers have throughput limited only by CPU and RAM, for which there's relatively little (though nonzero) overhead. IMO Python should have leaned into official adoption of gevent. It may not beat asyncio in raw performance numbers because asyncio can rely on custom-built bytecode instructions, whereas gevent has "userspace" code that must execute upon every yield. And, as with asyncio, you have to be careful about CPU-intensive code that may prevent you from yielding. But it's perfect for most horizontal-scaling soft-realtime web-style use cases.
- Buttons840 5y ago> our web servers and batch workers have throughput limited only by CPU and RAM Are you able to fully utilize a multicore processor? I'm not familiar with gevent.
- js2 5y agoThis is why you run one process per core, and you'll typically have something like nginx+uWSGI distribute requests across them. I use this combination with https://falconframework.org/ https://falconframework.org/ and boto3 to spool HTTP POST requests to S3 and SQS and am pretty happy with it. uWSGI supports gevent https://uwsgi-docs.readthedocs.io/en/latest/Gevent.html https://uwsgi-docs.readthedocs.io/en/latest/Gevent.html Falcon also benefits from Cython acceleration. It's been a while but at the time I tested against PyPy and either it was slower or had quirks I was unable to resolve (don't recall w/o consulting my notes).
- int_19h 5y agoHow would those green/virtual threads interface with native async APIs (e.g. the entirety of WinRT)?
- dapids 5y agoI still use greenthreads every chance I get. IMO The asyncio headaches are just not worth the abstraction hell involved in their event loop design, without being forced to stick to fragile concepts or consistently staying up to date with what is the current best way to do something with asyncio.
- birdyrooster 5y agoI don’t see what is so weird about it. The syntax is simple and straightforward. It’s how I wrote a ton of Python and it works great as long as you know what blocks and what doesn’t. Using aio_multiprocess it’s easy to saturate all the cores on a machine using the same basic syntax. It’s lovely really but still Python is too slow compared to golang so I rarely use it.
- andrewstuart 5y ago>>> In an alternate Python timeline, asyncio was not introduced into the Python standard library, I totally disagree with this negative perspective of asyncio. async programming is MUCH easier to program and understand than other concurrency solutions like multithreading. async works extremely well as a programming model as evidenced by the fact that it's exactly the model being adopted by all languages that want to implement practical and developer friendly concurrency. When I read such a negative take on asyncio I assume only this is a developer who hasn't done alot of programming with it, and is therefore still somewhat lacking in full understanding.
- alfanerd 5y agoYes, yes, yes and I hope the alternate timeline will start now. IMO the reason that we have so many ways of doing concurrency in Python (fra asyncio, curio, trio to gevent, multiprocessing etc.) is that it was never properly dealt with. It has to be built into the language, a feature of the language, like fx. garbage collection. The model of Pony or Erland, where are execution thread is started per core could also have been used in python, instead we got the async/await mess, which almost created a whole new language, where every library has to be rewritten. It saddens me to think that the ugly cludge that async/await is got added to Python almost without any discussion, whereas the insignificant walrus operator got so much heat that GVR quit.