10 ms·
An easy way to concurrency and parallelism with Python stdlib
- pdimitar 3y agoDoes not seem exactly like an easy way to me. Not super hard, surely, but not "easy". More like "moderately easy to do and a bit annoying to implement". Probably 20% of the effort shown in this post could have been expended to just write something very similar in Golang, and it would have taken less time, too. Because the way I see it this is trying to emulate futures / promises (and it looks like it's succeeding, at least on the surface). That can spiral out of comfortable maintainable code territory pretty quickly. But especially for something as trivial as a crawler, I don't see the appeal of Python. You got a good deal of languages with lower friction for doing parallel stuff nowadays (Golang, Elixir, Rust if you want to cry a bit, hell, even Lua has some parallel libraries nowadays, Zig, Nim...).
- simonw 3y agoIf you already know Python, the advice in this article is certainly a lot easier and more actionable than "just learn Go or Rust or Zig instead".
- pdimitar 3y agoCertainly. My point is that if you need to write that much code and/or do that much research, at one point the effort of doing it in another language will be less than to keep insisting on using a tool that's not designed for it. It happened with me and many other former colleagues. Though obviously, everyone decides for themselves when does that point come -- or if it comes at all.
- FreakLegion 3y agoThe point of the article is a handful of lines. The rest is accoutrement like the URL list and timing code. But sure, if tasks = {} for url in URLs: future = executor.submit(fetch_url, url) tasks[future] = url bothers you, this is perfectly (some would say more so even than the original) Pythonic: tasks = {executor.submit(fetch_url, url): url for url in URLs}
- throwaway46281 3y agoAs a side note, using a future as a map key struck be as a bit weird, though perfectly valid. It'd be more natural IMO to use a list for the futures, and have the fetch_url function return a (url, result) tuple. Or use the url as the map key and just iterate over the map items instead of using as_completed on the keys
- networked 3y agoI have found another way in the documentation for `concurrent.futures`. You can use `Executor.map` (https://docs.python.org/3/library/concurrent.futures.html#concurrent.futures.Executor https://docs.python.org/3/library/concurrent.futures.html#co...). It eliminates the need to wait on the futures explicitly. def main(): with ThreadPoolExecutor(max_workers=len(URLs)) as executor: for url, title in zip(URLs, executor.map(fetch_url, URLs)): print(f"URL: {url}\nTitle: {title}") The default value of `max_workers` since Python 3.8 has been min(32, os.cpu_count() + 4) You should probably avoid max_workers=len(items_to_process) It will not save memory or CPU time when you have few items (workers are created as necessary) and may waste memory when you have many.
- pid-1 3y ago> and/or do that much research, Is reading the official docs section on concurrency lots of research?
- Demiurge 3y agoWhat “much research” are you talking about? The amusing part is that the article calls out two groups of people into which your advice falls. It’s not that much code, it’s about 4 lines of code, creating a “pool” and calling a wait on future objects. This is a perfect solution for Python developers who have been perfectly happy using Django for years, and just need to scrape some API or download multiple files. No, they shouldn’t switch to a different language the moment they need to optimize something embarrassingly parallel, they can see whether a simple solution in stdlib is enough, and probably move on.
- oefrha 3y agoIf this is too much research for you, wait until you have to deal with the many problems of Go channels in the real world. (Reasonably well-known though controversial article: [1]) Don't even get me started on Rust. Concurrency and parallelism is hard. Yes, I've written a shit ton of code in all aforementioned languages. [1] https://www.jtolio.com/2016/03/go-channels-are-bad-and-you-should-feel-bad/ https://www.jtolio.com/2016/03/go-channels-are-bad-and-you-s...
- rich_sasha 3y agoPython is surprisingly bad at parallelism, for a data or framing workhorse. What TFA doesn't say is that process pools are quite fragile, certainly on Mac and Windows, but Linux also. They rely on pickling which is also fragile. That said, asyncio works surprisingly well if what you want is non-blocking execution and are happy with 1 cpu. But no parallel speed up.
- rmbyrro 3y agoAfter learning clojure, I found python's approach to concurrency terrible at best. Clojure is extremely easy to understand. It has basically three solutions, each for clear and defined use cases. It's much easier to judge what you should implement given a particular problem and how to do it. I wish Python had similar solutions.
- hleszek 3y agoThe easiest and modern way is simply to use asyncio...
- tonetheman 3y ago[dead]
- slig 3y agoIs there a way to add tasks with independent timeouts using only the Python stdlib? I was reading a piece of code yesterday that had `pebble` as dependency and it looked like it was only needed for the `pool.schedule(..., timeout=1)`.
- samsquire 3y agoThank you for the article. I use multiprocessing and I am looking forward to the GIL removal. I would really like library writers and parallelism experts to think on modelling computation in such a way that arbitrary programs - written in this notation - can be sped up without thinking about async or parallelism or low level synchronization primitives spreading throughout the codebase, increasing its cognitive load for everybody. If you're doing business programming and you're using python Threads or Processes directly, I think we're operating against the wrong level of abstraction because our tools are not sufficiently abstract enough. (it's not your error, it's just not ideal where our industry is at) I am not an expert but parallelism, coroutines, async is my hobby that I journal about all the time. I think a good approach to parallelism is to split you program into a tree dataflow and never synchronize. Shard everything. If I have a single integer value that I want to scale throughput of updates to it by × hardware threads in my multicore and SMT CPU, I can split the integer by that number and apply updates in parallel. (You have £1000 in a bank account and 8 hardware threads you split the account into 8 bank accounts and each store £125, then you can serve 8 transactions simultaneously at a time) Then periodically, those threads can post their value to another buffer (ringbuffer) and then a thread that services that ringbuffer can sum them all for a global view. This provides an eventually consistent view of an integer without slowing down throughput. Unfortunately multithreading becomes a distributed system and then you need consensus. I am working on barriers inspired by bulk synchronous parallel where you have parallel phases and synchronization phases and an async pipeline syntax (see my previous HN comments for notes on this async syntax) My goal would be that business logic can be parallelised without you needing to worry about synchronization.
- wongarsu 3y agoAt least in what I do, I find 80% of my parallelism needs covered by pool.map/pool.imap_unordered. Of the remaining 20%, 80% can mostly be solved by communicating through queues or channels (though admittedly this is smoother in Erlang or Rust than in Python). Of course that's not true for everything, and depending on the domain tree dataflows can also be great. I remember them being very popular in GPGPU tasks because synchronization is very costly there.
- tedivm 3y agoI know this article is all about the stdlib, but having built multiple multiprocess applications with python I eventually built a library, QuasiQueue to simplify the process. I've written a few applications with it already. https://github.com/tedivm/quasiqueue https://github.com/tedivm/quasiqueue
- cle 3y agoI recently have been doing--what should be--straightforward subprocess work in Python, and the experience is infuriatingly bad. There are so many options for launching subprocesses and communicating with them, and each one has different caveats and undocumented limitations, especially around edge cases like processes crashing, timing out, killing them, if they are stuck in native code outside of the VM, etc. For example, some high-level options include Popen, multiprocessing.Process, multiprocessing.Pool, futures.ProcessPoolExecutor, and huge frameworks like Ray. multiprocessing.Process includes some pickling magic and you can pick from multiprocessing.Pipe and multiprocessing.Queue, but you need to use either multiprocessing.connection.wait() or select.select() to read the process sentinel simultaneously in case the process crashes. Which one? Well connection.wait() will not be interrupted by an OS signal. It's unclear why I would ever use connection.wait() then, is there some tradeoff I don't know about? For my use cases, process reuse would have been nice to be able to reuse network connections and such (useful even for a single process). Then you're looking at either multiprocessing.Pool or futures.ProcessPoolExecutor. They're very similar, except some bug fixes have gone into futures.ProcessPoolExecutor but not multiprocessing.Pool because...??? For example, if your subprocess exits uncleanly, multiprocessing.Pool will just hang, whereas futures.ProcessPoolExecutor will raise a BrokenProcessPool and the pool will refuse to do any more work (both of these are unreasonable behaviors IMO). Timing out and forcibly killing the subprocess is its own adventure for each of these too. I don't care about a result anymore after some time period passes, and they may be stuck in C code so I just want to whack the process and move on, but that is not very trivial with these. What a nightmarish mess! So much for "There should be one--and preferably only one--obvious way to do it"...my God. (I probably got some details wrong in the above rant, because there are so many to keep track of...) My learning: there is no "easy way to [process] parallelism" in Python. There are many different ways to do it, and you need to know all the nuances of each and how they address your requirements to know whether you can reuse existing high-level impls or you need to write your own low-level impl.
- xeromal 3y agoComing from C#, I honestly HATE python's multiprocessing and multithreading. Hell, I hate it's async await. I learned recently that in one mode, it pipes the values across the process and this made it impossible to use when passing along large pandas dataframes. I'm sure half of it is just my own lack of knowledge with python's abilities but C# sure made it easier. lol
- smallerfish 3y agoMaybe I missed it, but how do the threads circumvent the GIL? > When a request is waiting on the network, another thread is executing. I'm guessing this is the meat, but what controls that? What other operations allow the GIL to switch to another thread?
- gpderetta 3y agoMy understanding is that the GIL is typically released around blocking operations. Aside for allowing actual concurrency for I/O heavy programs, it would be a trivial way to deadlock if it wasn't.
- ameliaquining 3y agoPython functions implemented in C can release the GIL when they're doing something that doesn't directly involve manipulating Python objects, and then re-acquire it when they're done: https://docs.python.org/3/c-api/init.html#thread-state-and-the-global-interpreter-lock https://docs.python.org/3/c-api/init.html#thread-state-and-t... All I/O functions in the standard library do this when blocked.
- hot_gril 3y agoThis is a far better explanation than the usual opaque "it's concurrent but not parallel" that I'd argue isn't even correct (cause two C calls on separate threads are running in parallel if they don't hold the GIL). Or "it's multithreading but not multiprocessing" which misses the point.
- Lukeisun 3y agoAwesome article, use it a lot in a python project at work and it's quite nice how simple it is. I'm trying to replicate the python code but in Rust and it is slightly slower, more than likely my fault though as I'm new to Rust.
- thisisauserid 3y agoWhen dinking around in Ipython you need to use a fork for the "multiprocessing" library called "multiprocess." Parallelism in a Notebook isn't for everyone, but how would these changes affect it?
- capital_guy 3y agoThis is a really nice little guide. Much thanks to the author. Sometimes you just need to hit a bunch of APIs independently and don't want to switch your entire architecture around to do so.
- akasakahakada 3y agoDon't see MPI. Can skip this article.
- crabbone 3y agoMPI is not in the "standard" library, or am I behind the moving fast and break things?
- Iwan-Zotow 3y agoStdlib is keyword
- potta_coffee 3y agoIf I need concurrency these days, I just write it in Golang. My primary use for Python was one off scripts for cloud management / automation tasks. Today I write maybe 70% Golang and 30% Python.
- cpach 3y agoI agree. The team behind Go has thought a lot about concurrency right from the start, and it really shows.
- potta_coffee 3y agoConcurrency in Go is just so easy and powerful.
- crabbone 3y ago> For those, Python actually comes with pretty decent tools: the pool executors. Delusion level: max. You have to be in a very, very bad place when this marginal improvement over absolute horror-show that bare Process offers seemed "pretty decent". Python doesn't have good tools for parallelism / concurrency. It doesn't have average tools. It doesn't have even bad tools. It has the worst. Though, unfortunately, it's not the only language in this category :(
- paulddraper 3y ago> It doesn't have even bad tools. It has the worst. > It's not the only language in this category Soo....not the worst? :) Or tied for it? What do you find difficult/wrong with pool executors? Also, you reference "Process", but FYI the article talks about multiple threads, not multiple processes.
- hot_gril 3y agoPool executors only solve one kind of use case. They aren't a general solution to concurrency+parallelism. And they're still the worst version of this pattern, because despite using multiple OS-level threads with all the associated overhead, the GIL prevents most of the real parallelism from happening. And if you want full parallelism, you have to use multiprocessing.Pool, which adds pickling overhead and incompatibility.
- crabbone 3y ago> Soo....not the worst? :) Yeah... I know, it's hard to imagine that there could be more than one worst. But, as I have to practice these things with my 4 year old, I become more patient with adults who don't get the concept too. Imagine you are in a class and the teacher gives everyone a pencil and a sheet of paper. Now, you want to find out who has the shortest pencil. All students compare their pencils and turns out that there are several pencils that are of the same exact length, and those are the shortest ones at the same time. So, more than one student has the shortest pencil. But it doesn't end there. Not all sets which define a "greater than" relationship are totally ordered. In such sets it's possible to have multiple different smallest elements. Trivially, in a set that's not ordered, every element is the smallest. > What do you find difficult/wrong with pool executors? Difficult? -- I don't know. Wrong? -- Well, it's pretty worthless... does it make it wrong? -- That's up to you to decide. The idea of threads is bad for many reasons: one in particular is of how exceptions in threads are handled. But this isn't unique to Python. Python just made a bad decision to use threads in the language that's supposed to be "safe". Python thread implementation craps its pants when dealing with many aspects of threads. For example, thread-local variables. Since threads are objects in Python, you'd expect local variables to be properties on those objects... well the mechanism to use them is just idiotic and nothing like you would expect. When it comes to interacting with "native" code from Python, you'd expect some interaction with Python's scheduler so that the native code can portion its own execution, allow Python to interrupt it etc. but there's nothing of the kind. Even though we haven't even gotten to the pools yet, pools, obviously, don't address any of the thread-related problems. If anything, they only amplify them. Specifically, the pool from concurrent package is worse than its relative from multiprocessing package because it uses "futures". The whole idea of "futures" is somehow broken in Python because of the neverending bugs related to deadlocking. It's been repeatedly "fixed", but every now and then deadlocks still happen. Here's the latest one I know of: https://bugs.python.org/issue46464 https://bugs.python.org/issue46464 . I've gone once down the rabbit hole of trying to make a native module work with Python threads... there's no good way to do it, but pools, be it from concurrent.futures or from multiprocessing are both very bad for many reasons. I was hoping to be able to give users an ability to control how parallel my native code is through the tools exposed by Python already, but that turned out to be such a disaster that I've given up on the idea. Python's thread wrappers are worthless for the native code that wants to actually execute concurrently -- they are only designed to execute Python code, non-concurrently. Like I already mentioned, Python has no infrastructure to communicate to the native code its scheduling decisions, no thread-safety in memory allocation, the code is overall poorly written (as in missing const, other imprecise typing, memory-inefficient data-structures)... there are no benefits to using that vs rolling your own. Only struggle with bad decisions.
- eachro 3y agoSo what is the consensus view on how to do parallelism in python if you just have something that is embarassingly parallel with no communication between processes necessary?
- WinLychee 3y agoif you have a task that is easy to split, make a python script that runs on a subset of the task, split into N subsets, and write one output per process? Once they all complete, join together the outputs. Maybe https://docs.dask.org/en/stable/ https://docs.dask.org/en/stable/ is a good start if you want a framework. I don't think there's a consensus, it depends on the problem.
- hot_gril 3y agoPeople here mention Pool, and I've seen it many times. It's this: https://docs.python.org/3/library/multiprocessing.html#introduction https://docs.python.org/3/library/multiprocessing.html#intro... from multiprocessing import Pool def f(x): return x*x if __name__ == '__main__': with Pool(5) as p: print(p.map(f, [1, 2, 3])) This forks out up to 5 processes. f(x) runs fully in parallel for each input. The inputs and outputs sent between processes via pickling.
- hot_gril 3y agoThe article shows how to use ThreadPoolExecutor, but that's not fully parallel. For that, you need multiprocessing.Pool, which is slightly easier to use anyway, unless your data happens to be non-pickle-able.