4 ms·
In case the author reads this, there is an error in the "How to Execute a Blocking I/O or CPU-bound Function in Asyncio?" [0] It reads: > The asyncio.to_threa
by mpeg 4y ago
In case the author reads this, there is an error in the "How to Execute a Blocking I/O or CPU-bound Function in Asyncio?" [0]
It reads:
> The asyncio.to_thread() function creates a ThreadPoolExecutor behind the scenes to execute blocking calls.
> As such, the asyncio.to_thread() function is only appropriate for IO-bound tasks.
It should say it's only appropriate for CPU-bound tasks.
[0]: https://superfastpython.com/python-asyncio/#How_to_Execute_a_Blocking_IO_or_CPU-bound_Function_in_Asyncio https://superfastpython.com/python-asyncio/#How_to_Execute_a...
- wrigby 4y agoI think the article is correct here, actually - if you need to run a CPU-bound task, you’ll need a ProcessPoolExecutor.
- danuker 4y agoI second this. Python has a Global Interpreter Lock, so only one Python instruction executes at a time. Async and threads only help you with I/O (where the OS affords waiting for multiple operations in parallel). To execute CPU-bound code in parallel you need multiple processes. I think this is a major disadvantage of Python, because processes are much costlier to spawn, and if you implement long-running workers to avoid frequent spawning, you have to incur serialization/deserialization costs, because shared memory support is very rudimentary (in essence, just fixed size numeric code).
- mpeg 4y agoIt's just awkwardly worded, I think. It probably should include a mention of ProcessPoolExecutor in the following bit where it explains the run_in_executor function to make more sense. Especially as the whole section talks about passing CPU-bound task to a thread pool, but then recommends not to use a thread pool. Though I think the GIL isn't really that big a deal for most CPU-bound code, with plenty of third party libraries and even within the python standard library a lot of the code is not native python, and therefore releases the GIL lock.
- mpeg 4y agoWell, that section is talking about CPU-bound tasks, and both threads and processes are valid choices for CPU-bound tasks, with different tradeoffs. I think there is often a lot of confusion around that because a lot of tutorials simply say to use threads for IO-bound tasks and processes for CPU-bound tasks, without going deeper into the differences. Threads can easily share memory which often increases complexity and requires locks to run safely, with processes you avoid locking (including the infamous Python GIL) but also have to pass around data which might hurt performance. Regardless, it still doesn't make sense to say threads are only appropriate for IO-bound tasks, the official documentation [0] seems to lean towards preferring threads for mixing IO and CPU-bound tasks, and processes for purely CPU-bound tasks. [0]: https://docs.python.org/3/library/concurrent.futures.html#threadpoolexecutor https://docs.python.org/3/library/concurrent.futures.html#th...
- wrigby 4y agoI think the section is addressing both CPU and IO-bound tasks: > How to Execute a Blocking I/O or CPU-bound Function in Asyncio? But to be precise, we should differentiate between Python’s interpreter threads and OS threads. In general, OS threads are a great way to parallelize CPU-bound tasks (with the issues of locking you mention), but Python’s interpreter threads are not (because of the GIL).