7 ms·
Parallel tasks in Python: concurrent.futures
- simonw 9y agoI used the Python 2 backport of concurrent.futures for a project recently (parallelizing calls to an external API) and it worked fantastically well. It's a really nice model for doing concurrent outbound I/O in a bunch of threads.
- orf 9y ago> posted in Jan. 2017 Now we have asyncio and awesome libraries like aiohttp[1] you can get a much, much higher throughput than you'd ever achieve with threads with less code. 1. http://aiohttp.readthedocs.io/ http://aiohttp.readthedocs.io/
- sametmax 9y agoPlus asyncio has futures too, and with run_in_executor(), you can await something in a thread/multiprocessing pool from inside the event loop transparently.
- greglindahl 9y agoYes, I use this extensively in a web crawler -- network I/O is done in a main thread that's cpu bound doing networking stuff, and cpu-burning tasks like html parsing are done in child threads done by run_in_executor. Without trying very hard I can get to a load of 3-4 during ordinary web crawling.
- miracle2k 9y agoI found this article extremely persuasive, and it matches my own experiences with asyncio: The performance gain might be there if most of what you do is waiting for a network response, but even a small amount of data processing will make your program CPU bound pretty quickly. http://techspot.zzzeek.org/2015/02/15/asynchronous-python-and-databases/ http://techspot.zzzeek.org/2015/02/15/asynchronous-python-an...
- zackelan 9y agoNote that it's not either/or - you can dispatch work from an event loop to a thread pool (or a process pool) with loop.run_in_executor [0], while loop.call_soon_threadsafe [1] can be used by worker threads to add callbacks to the event loop. This means that the "frontend" of a service can be asyncio, allowing it to support features like WebSockets that are non-trivial to support without aiohttp or a similiar asyncio-native HTTP server [2], while the "backend" of the service can be multi-threaded or multi-process for CPU-bound work. 0: https://docs.python.org/3/library/asyncio-eventloop.html#executor https://docs.python.org/3/library/asyncio-eventloop.html#exe... 1: https://docs.python.org/3/library/asyncio-eventloop.html#asyncio.AbstractEventLoop.call_soon_threadsafe https://docs.python.org/3/library/asyncio-eventloop.html#asy... 2: Flask-SocketIO, for example, requires that you use eventlet or gevent, which are the "legacy" ways of doing asynchronous IO: https://flask-socketio.readthedocs.io/en/latest/ https://flask-socketio.readthedocs.io/en/latest/
- deleted 9y ago[deleted]
- xtrapolate 9y ago> "even a small amount of data processing will make your program CPU bound pretty quickly" I don't know what "a small amount" means. You are bound by hardware (cores, hyper-threading). It simply makes no sense to spawn 32 threads executing computation intensive code (no IO), on a machine with 2 cores.
- trentnelson 9y agoYou can get the best of both worlds with PyParallel! Async I/O and multiple threads. (Experimental project, not intended for production use, so I say this some what facetiously.) http://pyparallel.org/ http://pyparallel.org/
- Rotareti 9y agoLast commit Oct 2016. Not sure where this is going.
- mixmastamyk 9y agoCool, just wrote my first code with this module a week ago. A client needed to run background tasks under Flask without the ops complexity or dev time needed to set up a job queue. https://stackoverflow.com/a/39008301/450917 https://stackoverflow.com/a/39008301/450917
- pletnes 9y agoOne common misconception (or should I say, overgeneralization) is repeated in the article: threads are always unsuited to CPU intensive work. For instance, most numpy operations release the GIL, meaning that you can perform heavy computation on multiple threads simultaneously. Certain other C extensions do the same, including some bits of the standard library. The usual caveats apply about threading bugs, of course. Another detail is that numpy linked to e.g Intel MKL will multithread some operations by default. Running multiple threads-in-threads is likely to cause slowdown.
- jzwinck 9y agoOnly np.dot() has intrinsic multithreading. No other functions do. Bizzarely np.dot() is the fastest way to do things other than dot product (like copy or multiply) in some cases.
- vladf 9y agowell, also np.linalg routines that call LAPACK may be multithreaded.
- jzwinck 9y agoThat's misleading--you say linalg routines "may be" multithreaded, but the vast majority of them never have been. matmul and einsum, despite being candidates for intrinsic multithreading, are not multithreaded. You can read discussion about that here: https://jackkamm.github.io/blog/a-parallel-einsum/ https://jackkamm.github.io/blog/a-parallel-einsum/
- vladf 9y agoI'm sorry, isn't "objects belonging to class C may have property P" a fair way to say there are some members which have it and some which do not? I don't see how that's misleading. I was correcting your statement about np.dot being the only parallel fn
- 9y ago
- rcthompson 9y agoWill this finally let me write a parallel Python script that doesn't explode when I press control+C?
- metalliqaz 9y agoDepends. What do you want to happen when you press ctrl+C?
- rcthompson 9y agoI want it to not hang forever, not spam the console with pages of exception tracebacks, and not require a ton of boilerplate process-management code to accomplish the above. Ideally it would also allow me to handle other exceptions (i.e. besides KeyboardInterrupt) that occur both in the main process and child processes. I've never figured out how to do this with multiprocessing, despite lots of attempts.
- mixmastamyk 9y agoDon't know what you did wrong, but you should be able to catch exceptions at their appropriate level and shut down gracefully.
- icegreentea2 9y agomultiprocessing was designed to mirror the threading library... so the exceptions are specifically set up not to cross thread/process boundaries. But the same strategies to handle multithreaded exceptions and keyboardinterrupt should apply. If you want the child processes to shutdown when the main process goes down, you should be able to just set .daemon = True on the process objects before you start them. If you want exceptions in the children thread to propagate up to the main thread and then handle, looks like you'd just need to send the exception across a queue or something in multiproccessing. In the new futures library, your future (wrapping the process) has a return value property you can try to access. If the child process ran into an exception, that exception would be raised in the main process when you tried to access that return value.
- dokument 9y agoWould this allow separate GC'ng for each task?
- loganekz 9y agoOnly ProcessPoolExecutor[1]. If you use a thread pool or async io it will be single python process/GIL. [1] -https://docs.python.org/3/library/concurrent.futures.html#processpoolexecutor https://docs.python.org/3/library/concurrent.futures.html#pr...
- rflrob 9y agoWhat's the advantage of using a ProcessPoolExecutor over just using multiprocessing? Is it that there's a single interface that you can use for both threads and processes?
- rcthompson 9y agomultiprocessing also has a single interface for both threads and processes, so it's not that.
- icegreentea2 9y agoI don't really think there's an advantage. Just like how multiprocessing tries to mirror the threading interface, ProcessPoolExecutor just mirrors the threaded implementation of the new futures based concurrency interface. I think futures are nicer for certain types of interactions. For example, futures 'return' actual values, so its nice for dispatching a task that you'll get a result from back. Futures also raise exceptions (when you try to inspect their results, if an exception occurred in the task). This might make for cleaner error handling code.
- solotronics 9y agoso how does this compare to something like Deco? https://github.com/alex-sherman/deco https://github.com/alex-sherman/deco. I guess since this uses a single GIL its good for IO limited things?
- icegreentea2 9y agoAs written, the code in the blogpost is good for IO limited things. But as it notes, if you replace 'ThreadPoolExecutor' with 'ProcessPoolExecutor' then you get actual multiprocessing, and you may be able to get speedup on compute bound tasks. The linked repo looks like some nice wrappers/decorators around the 'old' multiprocessing library to make it really easy to parallelize a bunch of function calls within a blocking function.
- ggm 9y agoI tried using threads on multiple pipe reading, to centralise a logfile sorting problem (each discrete logfile is a gz which itself is only partially in order, and then between files a merge-sort has to be performed) It was enjoyable to try to fix, but ultimately I found the solution only marginally better than explicit processes feeding a single reader doing round-robin. I think the lesson I learned is that if the problem integrates back into a single context there isn't much you can do to avoid that bottleneck once all the other parallelism opportunities have been overcome.
- Jsharm 9y agoDoes anyone have a recommendation for what to use for a cache shared between processes? Would hdf5 work?
- cranklin 9y agoIf you want a file-based cache, yes.
- tyu100 9y agoI'm using concurrent.futures in production and its use of the multiprocessing module caused the Python grpc library to break in a really strange and hard-to-debug way: https://github.com/grpc/grpc/issues/13873 https://github.com/grpc/grpc/issues/13873 I suspect it's not the only Python library that will see issues if you are running it in the Future context.
- amelius 9y agoSad to see that Python still suffers from the Global Interpreter Lock (GIL), and that the only way out is still to use multiple processes (which causes problems of its own, e.g. sharing of large data structures becomes expensive).
- dullgiulio 9y agoOnly for computationally expensive operations done in interpreted Python. C extensions, IO operations etc. always release the lock. In practice GIL is a problem only when it is profiled to be a problem. Python is used a lot in the data analysis world and nobody cares about the lock, because a fraction of the CPU time is spent within the lock.
- amelius 9y ago> Only for computationally expensive operations done in interpreted Python. C extensions, IO operations etc. always release the lock. So you are saying it is a problem in Python but not in other languages? Which is exactly my point :)
- Rotareti 9y agoI wrote a similar interface to run asyncio compatible ProcessPool/ThreadPool executors, a couple of days ago: https://github.com/feluxe/aioexec https://github.com/feluxe/aioexec
- delaaxe 9y agoThe last for loop of the last code example doesn't need to be under the with statement.
- Dowwie 9y agoI found a concurrent.futures.ThreadPoolExecutor useful for database seeding, where I invoke a whole lot of sql alchemy core inserts