10 ms·
Why Not Python - the GIL hinders concurrency
- kevingadd 13y agoThis feels to me more like a critique of fork() and the unixy process-oriented parallelism model than a critique of Python. Of course, the author mentions this as a caveat, but it makes me wonder if the blame is being laid where it should be (and also whether in some scenarios like this, you should really just build applications the way applications are normally built for a platform, no matter how much you dislike it)
- willvarfar 13y agoMultiprocessing is touted to work around the GIL. I too have run into the wall trying to get big problems solved in Python. And when I went multiprocessing, I ran into limitations in its internals e.g. using `select()` that really surprised me too.
- zhemao 13y agoThe reason why Python gets the blame is that in languages without a GIL, you can choose to use threads instead of processes in order to get better performance. Python's GIL more or less restricts you to using processes for CPU-bound concurrency. Also, as others have mentioned. Two processes can share memory (at least on Linux) by explicitly using the shared memory system calls. However, this is complicated in Python by the need for reference counting of objects created in the shared memory region.
- MostAwesomeDude 13y agoThe author appears to not be aware of where the big leagues are. Additionally, as is becoming a recurring theme here, he doesn't know about PyPy nor Twisted. This is continually disappointing.
- dbecker 13y agoHe may know about PyPy, but need to use libraries that only work with CPython. Numpy is an example of this type of library, though there are many others. Unfortunately, PyPy isn't a viable alternative for most scientific computing problems...
- hcarvalhoalves 13y agoI would say PyPy isn't a viable alternative for any problems at this point, given the general lack of documentation.
- gsnedders 13y agoWhat documentation does it lack? What are you looking for to be documented?
- hcarvalhoalves 13y agoLast time I looked into PyPy (2-3 months ago): - There wasn't up to date documentation about basic topics, like how to compile it. - There isn't comprehensive documentation about what RPython is supposed to be. - In the repository, it's hard to figure out what is RPython and what isn't, what's code for the interpreter, what's code for the stdlib, etc. Specially because modules import from each other in crazy ways. - It's hard for someone who's not a contributor to peek at the code to figure out why client code isn't running on PyPy, since the only people who understand the architecture are the authors. - The code itself is pretty opaque and light on comments. I understand it's a fast moving project and that it had major rewrites so far, but those are the reasons why I say it's not a viable alternative for production.
- cdavid 13y agoIf you think twisted is a solution to the problems mentioned in the OP, you haven't understood the problem. Twisted may be a solution to IO bound processes, where you do cooperative parallelism instead of preemptive (aka threads). It is utterly useless for CPU bound processes (e.g.: you want to compute some expensive operation on top of a big numpy array, twisted does not help you with that at all).
- lmm 13y agoTwisted would work perfectly for the message passing example given in the article.
- coldtea 13y ago>The author appears to not be aware of where the big leagues are. Additionally, as is becoming a recurring theme here, he doesn't know about PyPy nor Twisted. This is continually disappointing. Do you even know the author? I'm pretty fucking sure he does know about PyPy and Twisted. He is a HN regular, a good tech blogger, and he has TONS of experience with Python for production use.
- AnIrishDuck 13y agoAll of these problems can be addressed using inter-process shared memory. Shared memory support is built into multiprocessing [1]. Now I agree it might not be convenient, but that's a matter of libraries. This post would be much more constructive if it was speculation on what such a library could look like, instead of pretending that concurrency "doesn't work" in single-threaded Unix processes. 1. http://docs.python.org/2/library/multiprocessing.html#module-multiprocessing.sharedctypes http://docs.python.org/2/library/multiprocessing.html#module...
- yummyfajitas 13y agoOk - using shared ctypes, how would you build the event processing system or the cache described in the blogpost? I'm sure it's possible to build an IPC garbage collector and enable everything described. But to describe it as "not convenient" is perhaps a small understatement.
- AnIrishDuck 13y agoWell, it sounds like a good start - at least for the cache example - would be a shared memory hash table. Building one in pure python with the existing multiprocessing primitives is doable, though obviously not trivial. Building one in C is also possible, though obviously there's the development time / performance tradeoff there. A shared memory hash table actually seems like a really useful thing to eventually make its way into the multiprocessing module. For the event example, parsing the data structures into shared memory and then accessing from there would remove the parsing and memory overhead he complains about. I don't really want to dig in to the standards he cites and see how difficult that solution would be, but again, I can't see a reason why it isn't possible. At any rate, in both cases, the solution is definitely doable; there's nothing inherent in python's concurrency model that "prevents you from properly making use of modern multicore hardware" Side note: shared memory access in modern multicore hardware may not get you the performance gains you would initially think. Especially for write-heavy workloads, cache invalidation is a huge problem.
- yummyfajitas 13y ago
- senko 13y agoTimely and related talk given today at EuroPython conference about concurrency in Python and why (and when) GIL doesn't matter: http://www.youtube.com/watch?v=b9vTUZYmtiE http://www.youtube.com/watch?v=b9vTUZYmtiE
- ditados 13y agono mention of gevent, celery, etc. As someone who runs thousands of concurrent tasks in a mix of process/gevent (one UNIX process for each 100 greenlets across 48 cores on two boxes), I find the OP's toliling rather misguided.
- zhemao 13y agoThe author's use case is clearly very different from yours. He is talking about CPU-bound processes which need to share a large amount of memory with each other. In this case, multiprocessing and message passing is not really the best fit. Multithreading or shared memory results in far less CPU usage and memory duplication.
- yason 13y agoThe standard Python workaround to the GIL is multiprocessing. Multiprocessing is basically a library which spins up a distributed system running locally - it forks your process, and runs workers in the forks. The parent process then communicates with the child processes via unix pipes, TCP, or some such method, allowing multiple cores to be used. Multiprocessing is the way to do parallelism. Deviating from that should be an exception -- for example, shared memory maps could be used to transfer select data objects instantaneously between the processes instead of serializing/deserializing over a pipe, and only those while still retaining separate process images. Threads were practically invented as a compensation for systems with heavy process image overhead. I think Python is very Unix in this regard. And that's not a bad thing per se. Unix and Linux can do multiprocessing very efficiently.
- dakimov 13y agoOrly? If your language is handicapped maybe. In normal languages you have no problems with threads, and the assertion that multiprocessing is the way makes no sense.
- zbowling 13y ago> Multiprocessing is the way to do parallelism. Deviating from that should be an exception Unless you have a language that doesn't break down when you use threads. Threads are unquestionably easier and better to use for multiprocessing when they work without the tools you are using breaking that.
- yummyfajitas 13y ago
- andrewguenther 13y agoGIL doesn't hinder concurrency, Python's threading library is still concurrent. It just isn't parallel.
- zbowling 13y agoWhat? The entire definition of being concurrent is doing work in parallel.
- jamesmiller5 13y agoI believe that is a common misconception. Concurrency enables parallelism but it isn't a requirement. The golang community has lectured about this at length: http://blog.golang.org/concurrency-is-not-parallelism http://blog.golang.org/concurrency-is-not-parallelism
- waxjar 13y agoIt's not :) Running things concurrently means that you finish two or more tasks in the same time period. Running things in parallel means those things happen at the same time.
- coldtea 13y agoNope. Concurrent can be interleaved (like you get multitasking in a single CPU). Parallel is actually parallel (same time).
- keeperofdakeys 13y agoConcurrency deals with multiple threads/processors running together, but not necessarily executing at once. So it applies everywhere, even on a single core machine. A great example of concurrency is a gui application. Without threads, any blocking operations will make your application unresponsive. You need multiple threads so that if one thread is blocking on a syscall, the other is available to execute gui actions. This works in python, because the GIL only applies to actually executing python threads. As far as writing general-purpose applications in python goes, most don't hammer all cores with work. In fact, they usually only require I/O concurrency, so python works as well as any other language.
- 3amOpsGuy 13y agoHow can the GIL, which is restricted to one process, hinder concurrency? It can only impact 1 single form of concurrency, threading. Why use threads? Noone does parallel compute on CPUs these days, not since GPGPUs rocked up almost 5 years ago (and we often use python as the host language, thanks pyCuda!) Parallel IO then? Well, except that async IO is often far more resource efficient (at the cost of complexity though). Threading is dead(-ish) because its hard to write, hard to test and expensive to get right. Concurrency in python is very much alive though.
- zbowling 13y ago> Noone does parallel compute on CPUs these days, not since GPGPUs rocked up almost 5 years ago I want to live in your world where all you are processing is vectors and FFTs in parallel on GPUs and not doing real work (accessing databases, processing data from sockets, etc). Threading is not dead. It's only crippled in python so everyone wants to invent ways of saying it is dead. Threading being hard to write is also a fallacy. I use thread backed dispatch queues which make concurrency simple in my language of choice right now. Threading like that is easy thanks to closures and a good design patterns. My apps are entirely async and run heavily parallel and it's easy to maintain and write using that.
- 3amOpsGuy 13y agoAccessing databases, processing data from sockets, are not CPU bound activities? I believe you've misread my post. For all your IO cases, and all your cases are IO, would you, and future maintainers of your code, not be better served with simpler abstractions which permit scaling past a single host?
- zbowling 13y agoI wasn't referring to the IO bound side of it but the general work involved with everyday generic work that was not something that a GPU can do very well. It's silly to say the answer to doing parallel is to throw it on the GPU. But referring to the IO side debate, the current design of many of the libraries that you call in the C world are often inheirtly blocking. 'gethostname' for example is a blocking call. There is no async version of it. To use them without contention on your single threaded application, you have to call them from worker threads. The common pattern is to spin up a thread to call it and do work on it. It's easier often to have your workers be thread bound like that to simplify your code and only lock shared resources when you need them. I can also make a massively async version of all my code that handles everything using async methods and in many cases this is better but it's harder to write and not always an option. Something I have to deal with daily because I run into the C10K problem all the time at work (http://en.wikipedia.org/wiki/C10k_problem http://en.wikipedia.org/wiki/C10k_problem). Even in the async model though I still want to be running code in parallel and I would still rather build that model up with thread powering it and not multiple processes and shared memory.
- pdpi 13y agoThe guys at CCP (the makers of EVE Online) seem to be doing just fine with parallelising stuff in Python.
- zbowling 13y agoNo one said you can't be parallel in python. The problem is that you can't use simple threads and must resort to separate processes, shared memory, and IPC to shard out your work that way. CCP uses twisted which manages to help you with yielding when doing async IO keeping as much work off the GIL when you are waiting on data and sockets and builds in cooperative multitasking concepts to let you yield to other work, but it's internally not multithreaded or multiprocessing out of the box. You still have scale up worker processes in some cases (usually one per CPU you have) to really make it effective.
- zzzeek 13y ago> Memory duplication has a relatively simple solution, namely using external cache such as redis. But the thundering herds problem remains. At time t=0, each process receives a request for f(new input). Each process looks in the cache, finds it empty, and begins computing f(new input). As a result every single process is blocked. this is incorrect. The processes coordinate on a lock held in redis itself. This solution is available right now using the Redis backend in dogpile.cache (of which I am the author): https://dogpilecache.readthedocs.org/en/latest/usage.html https://dogpilecache.readthedocs.org/en/latest/usage.html
- yummyfajitas 13y agoI was unaware that redis had that feature, I'll update the blog post to reference it. Thanks.
- njbooher 13y agoIt's also not too hard to DIY: http://www.dr-josiah.com/2012/01/creating-lock-with-redis.html http://www.dr-josiah.com/2012/01/creating-lock-with-redis.ht... https://github.com/njbooher/boglab_tools/blob/dece35f13a8fcb9cb5b7eefdee2d6f9916350918/entrez_cache.py#L105 https://github.com/njbooher/boglab_tools/blob/dece35f13a8fcb... When multiple processes ask this for the same file one of them downloads it and the others wait for it to finish.
- thezilch 13y ago> The implementation would also likely be considerably more complicated than the 160 linues of code that the Spray Cache uses. Not likely, using Twisted deferreds and a sane cache-wrapper with herd awareness -- you probably want this regardless of long-running cachables. Of course, Python 3.2 also has a futures [thread or process] builtin, if that's your thing. 10-20 lines of code.
- cdavid 13y agoNumPy developer here. First, I agree that the "the GIL is not an issue" is an annoying meme. It can be an issue, and when it is, it is annoying as it complicates some architectures. I don't think those architectures are as common as people usually think. A few remarks: I would qualify the GIL as a tradeoff rather than a mistake. You loose CPU-bound parallelism with threading, but you gain easy to write C extensions (well, relatively speaking). I think this point is critical to the existence of something like NumPy (which is unrivalled in general PL AFAIK). If you need to share a lot of data (a big numpy array), then you can use mmap arrays, there is no serializing/desarializing cost, and that's efficient. See for example some presentations by O. Grisel from scikits-learn (https://www.youtube.com/user/EnthoughtMedia/ https://www.youtube.com/user/EnthoughtMedia/ -- disclaimer, I work for Enthought, but all those video happened at scipy 2013). IMO, the only significant niche where this is an issue is reaching peak performance on a single machine, but python is not really appropriate anyway -- that's where you need fortran/C/C++, and even scala/java/haskell are quite unheard off there.
- ProblemFactory 13y agoThis is an interesting point. The existence of GIL may even indirectly speed up numerical Python code. Numpy and cython are significantly faster than N pure python threads, and the GIL encourages their development.
- cdavid 13y agoI would not go as far as saying the GIL speeds things up :) Everything else being equal, I would rather not have a GIL. I don't know any efficient runtime which uses reference counting and has no GIL, and integrating C with garbage collection is generally hard. That's one of the main reason why integrating C/etc.. with JNI or Matlab is a nightmare. The only one that managed to reduce the impedence mismatch and that I am aware of is Lua.
- falcolas 13y agoGiven that early efforts to remove the GIL in cPython resulted in slower single-threaded execution, I would say that yes, the GIL does speed some things up.
- halayli 13y agoif you need performance, just don't use Python. You'll be better off using C/C++.
- pjscott 13y agoCython is also definitely worth looking at, if you want to write Python code and you need parts of it to be fast. http://cython.org/ http://cython.org/
- halayli 13y agoIf it's a small part you want to optimize it works great. But if all your project is performance sensitive it becomes a big hack that's hard to maintain.
- VeejayRampay 13y agos/Python/Ruby. Title still works.
- shanemhansen 13y agoHe has a valid point. There are certain types of workloads which don't scale well on a single machine when writing pure python code. If your workloads happens to need to get more cpu performance out of a single machine in pure python and isn't easily to parallelize using processes, python's not the best choice. Personally I feel like go is good for this use case because I consider writing my own object lifecycle code a pita. Others will only use multithreaded c code, something I'm not eager to touch. I do however feel the need to make a couple clarifications: If you fork a process, you don't necessarily duplicate memory. Yay for COW. Fork twice and you've got fanout with very little extra memory. Threading works fantastically for I/O. C libraries can release the GIL during processing. So if you have to do something computationally expensive, you can let another python thread run. I rarely feel restricted by python, and when I do I usually find that other aspects make up for it. I find it makes a decent systems programming language (although not as good as C). The os library and the ctypes library give you the capability to do most (all?) system level tasks. I tend to prefer go over python when library support isn't an issue, but that's more due to it's type system and channel primitives. As someone who's pushed python to it's limits as a systems programming language I'm pretty comfortable dropping down to C isn't often called for, and writing in java/scala isn't ever required.
- cdavid 13y agoCOW semantics of fork is not that useful with python because of reference counting (at least cpython implementation where the reference count lives inside the object memory representation). It may work much better with pypy (which uses a 'real' GC but still has a GIL)
- shanemhansen 13y agoExcellent point. Ruby's take on this has been interesting. http://patshaughnessy.net/2012/3/23/why-you-should-be-excited-about-garbage-collection-in-ruby-2-0 http://patshaughnessy.net/2012/3/23/why-you-should-be-excite... I'm not a ruby expert but by understanding is that basically they've moved the refcount field out of the struct and out of the memory page. It would be nice if python did something like this. [edit] My summary of what ruby does is totally wrong and while this is interesting and applies to memory managment, doesn't necessarily apply to refcounting.
- falcolas 13y agoThis is just my opinion, but it's served me well in the 8 or so years I've done Python development: Why are you doing compute intensive work in Python? Python is not well optimized for doing compute intensive work. It's made some strides over the years, being able to do limited bytecode optimization and multiprocessing, but it's not a high performance computation language. However, it is a fantastic glue language. Write your compute intensive portions in C, and use Python to glue the portions of C together. You can even do your thread generation in Python, release the GIL during your C execution, and you don't even have to worry about the complicated process of spinning up threads in your C. In other words, use the language like it was designed to be used.
- cgh 13y agoNo offense, but I don't think you understand what numpy and scipy are. Numpy is the specific use case mentioned in the post.
- falcolas 13y agoI know precisely what Numpy and scipy are. Their choice to not release the GIL during their operations is not necessairly a limitation of Python. Of course, not releasing the GIL eases interaction with Python level variables, so that might be a tradeoff they decided is worth losing threading for. But please don't blame Python for the choices made by a library developer.
- pwang 13y ago> Their choice to not release the GIL during their operations is not necessairly a limitation of Python. Um, what? They do not hold on to the GIL when doing compute intensive operations.
- bad_user 13y ago> Write your compute intensive portions in C This advice is so often given by people, that I suspect that most people don't know what they are talking about, because in fact most people never do that, being unable to drop to C. I dare you to show me a piece of "compute intensive" portion that you optimized with C. I've worked with Python for about 3 years. After struggling with it to stretch the boundaries of what it could do, I finally gave up in frustration and the final solution was to find a less limiting environment. Worked great thus far and I don't miss anything about Python.
- minimax 13y agoI don't get how having multiple copies of a quote hanging around in memory is a big problem. It's probably less than 100 bytes total (symbol + side + price) and probably significantly smaller than the size of the state associated with the "statistics process."
- yummyfajitas 13y agoAn individual quote isn't a problem, but they don't tend to come one at a time. For realtime processing, serialization costs are usually the biggest issue, as is cache locality. On the other hand, for batch processing, memory can be a big issue (CPU less so).
- binarycrusader 13y agoBut fundamentally, the GIL prevents Python from being used as a systems language. This is only true when severely restricting the definition of a systems language. The vast majority of command line utilities you'll find in a typical operating system are often primarily I/O-bound or do not have critical performance requirements. As such, I'd only be willing to agree with the author's statement for a very specific subset of systems programming. For example, I've worked on a packaging system written in Python for the last five years or so. The package system is primarily I/O-bound the vast majority of the time (waiting on disks or network), and almost all significant performance wins (some as much as 80-90%) have come from algorithmic improvements rather than rewriting portions in C (very little is written in C). As one of my colleagues is fond of saying (paraphrasing), "doing less work is the best way to go faster". It also ignores the fact that depending on the problem space involved, there may be readily available solutions that provide excellent performance that don't involve threading (e.g. the multiprocessing module, shared memory, etc.).
- Nursie 13y agoRegardless of other arguments, it basically means that threads in python become only a way of thinking about a problem rather than a way yo utilise hardware fully. It's a shame.
- pwang 13y agoNice post, Chris, and it's definitely a problem that many people are confused about (as the comments in this thread show). People who think the GIL is a problem either have no idea what they're talking about, or they really know what they're talking about. People who don't think the GIL is a problem either have no idea what they're talking about, or they really know what they're talking about. :-p One of the lesser-appreciated facts about the GIL is that it is an implementation detail of CPython. That is, it is entirely possible to implement a Python that does not have process-level globals and C statics. There is no structural reason why a C or C++ implementation of Python needs to have a GIL. It's just a legacy of the implementation that has stayed around for a long, long time. For instance, Trent Nelson has done some work to show that you can move to thread-local storage for the free lists and whatnot, and get nice multicore scaling, even with the existing CPython codebase[1]. There are still concurrency primitives that the language would need to offer at the C API level to manage the merger of these thread-local objects, but it's a whole lot better than only being able to use a single core in modern days. Fortunately I mostly get to work in the scientific field with (mostly) data parallel problems. [1] https://speakerdeck.com/trent/parallelizing-the-python-interpreter-an-alternate-approach-to-async https://speakerdeck.com/trent/parallelizing-the-python-inter...