16 ms·
Removing Python's GIL: The Gilectomy [video]
- kris-s 10y agoI hope this is a fruitful endeavor but I can't help but feeling the pressure to "Not have another Py2 -> 3" situation will win out and the GIL will remain. Which, I think is completely fine. For me Python is the ultimate glue code - it's the duck tape of my programming world. If I know multi-core performance is going to be an issue up front, I would pick another language.
- smegel 10y agoIdeally Golang. No other language really does it right like Go - unifying event based and traditional multi-threading paradigms in a way that transparently utilizes all the cores on your system, while allowing you to write plain old iterative, blocking code. Go may be less than ideal in many other regards (i.e. the rest of the language), but it gets this right.
- rubiquity 10y ago> No other language really does it right like Go Not trying to start a war but... Erlang, Haskell, F#, and a few others abstracted Evented IO and parallelism before Go even existed.
- smegel 10y agoI should have added mainstream, non-functional, compiled language.
- InclinedPlane 10y agoRust? I would never consider Go mainstream myself though. Erlang is practically more mainstream.
- smegel 10y agoRust may well be, but it is not out yet? Erlang is mainstream, but being functional probably not open to consideration for 99% of developers.
- dalke 10y agoRust, version 1.0.0, was released in May 2015, so over a year ago. 1.9 came out 10 days ago. Is that not out enough for you?
- smegel 10y agoI have no idea. Is the language spec fully implemented by a robust and stable compiler/runtime? If so, I'll use it today!
- deleted 10y ago[deleted]
- geofft 10y agoYes, the compiler has fully implemented everything the language is intended to do since the 1.0 release a bit over a year ago. (Work focuses on the compiler and standard library itself; there isn't a separate spec doc that is updated to do things that aren't yet implemented or being implemented, so it's more like Python or GHC than R6RS or C++.) The compiler and the language are both stable; there's been no "runtime" in a strict sense since a few months 1.0, but there's a standard library which is usually statically linked in, and both the language and the standard library are covered by a very strong stability guarantee (more-or-less, code that works today and isn't deliberately doing something stupid will work indefinitely, as there are no plans for a Rust 2.0). Plenty of people are using it in production. So I'd go with "Yes."
- foota 10y ago@smegel, yes it is.
- nzjrs 10y ago
- whateveracct 10y agoMoving the goalposts a bit there, aren't you?
- smegel 10y agoNo, I think they are common assumptions.
- chrisseaton 10y agoThey obviously aren't - this whole thread is about Python, which is not a compiled language. So clearly nobody moving from Python to Go already had a requirement for it to be compiled.
- smegel 10y agoI was talking about the points "mainstream, non-functional" which excluded all the previous counter-examples. I only added "compiled" as it is a notable feature of Go that may be seen as beneficial for performance and other reasons.
- rdtsc 10y agoIf Erlang breaks there is a 50% you won't be able to see cat pictures on your smart phone anymore. It might not be cool with the startup crowd or the Google-following one, but is a an industrially used langauge for many products (RabbitMQ, CouchDB, Riak, Ejabberd, WhatsApp messaging, payment systems in Europe, and others). It is cool you like Go, it is a nice language. But before you claim "it is the only language that does X", you should learn about other langues as well.
- smegel 10y agoI know about Erlang and what it can do, and I maintain Go is the only mainstream, non-functional language that does abstraction of event based/threaded programming well, or at all. I'm not even sure Rust does...the best I can find is it was planned to, but it didn't make it in. And I don't consider any functional language as a viable choice for the vast majority of mainstream users.
- deleted 10y ago[deleted]
- rubiquity 10y agoOutside of HackerNews and Google I wouldn't consider Go to be very mainstream at all.
- Pxtl 10y agoEven good old c# with async/await does async I/O pretty well.
- ngrilly 10y agoC# with async/await is a different approach: it's a kind of cooperative multitasking where "yield points" are explicit. Go uses a kind of preemptive multitasking, in which goroutines don't have to explicitly "yield" control to the scheduler.
- fatbird 10y agoHastings proposed avoiding the breaking changes issues by having separate extension integration points for python with the GIL and without. In other words, cpython would be compilable without the GIL, and extensions wanting true threaded parallelism would implement additional integration points to take advantage of that if run in a GIL-less python. So numpy, for example, would continue to work with the GILled python, but a data scientist would also have the GIL-less python available, set up a virtualenv for it, install numpy, and see all those benefits, without ever causing breaking changes for extension authors. Dunno if that'll work in practice, but it seems like the best possible plan going forward to get the change without causing strife.
- IndianAstronaut 10y ago>If I know multi-core performance is going to be an issue up front, I would pick another language. This is why I have shifted most of my code (I work in data warehousing) to Go. Python is handy for little scripts and data mining.
- IndianAstronaut 10y ago>If I know multi-core performance is going to be an issue up front, I would pick another language. This is why I have shifted most of my code (I work in data warehousing) to Go. Python is handy for little scripts and data mining.
- dbcurtis 10y agoThis was a great talk, and was presented to a packed house (see below). If you don't want to watch the whole video, the issue really boils down to this: 1) Lot's and lot's of fine grained locks. One on every collection. Oooof. 2) All that locking and unlocking absolutely thrashes the cache with Larry's current prototype implementation. I was surprised at how little time was spent doing the actual locking/unlocking itself. I was blown away by the performance impact of maintaining cache coherent locks (at least in Larry's current implementation). And for reference, I was a logic designer on mainframes in the 1980's where we paid attention to making locks perform well across independent caches, so I'm no newb around this issue, but it was still striking to me. (packed house) Larry's joke at the beginning about "practicing for getting on your plane later tonight" is a reference to the packed seating. The Portland convention center staff were taking the Portland fire marshall's directives quite seriously. We spent several minutes making sure every seat was occupied, and then the staff evicted the standee's :/
- chrisseaton 10y agoI wonder if you could completely elide the CAS needed by userspace locks (the bit which thrashes the cache I presume) while there is only one thread running. Then when a new thread is created for the first time (if it ever is), pause the program and promote and acquire all the locks. Like how the JVM assumes classes are final and removes virtual calls until it first sees a subclass.
- fatbird 10y agoLate in the talk he discusses strategies to avoid thrashing that amount to making locks as local as possible, ideally such that they behave the same way they would in the current unsynchronized environment. The big problem is that all the locks are global, so a change to a lock triggers a gigantic cascade of cache invalidation (and voided branch prediction, etc). It's a great talk; the video is really worth watching. At the end, he says he welcomes anyone to join him in the sprints, but basically don't bother unless you really know cpython and multithreading programming because this is very advanced stuff.
- Bromskloss 10y agoIs there anything inherent to the language that makes this global lock thing difficult to be without in Python, but not in other languages?
- dalke 10y agoInherent? No. Neither Jython nor Iron Python have a GIL.
- jerf 10y agoIt was ground in to CPython at an early date and every Python C extension ends up with it ground in, too, because it's in the C API. Some things are very difficult to retrofit on to an existing system. [1] If you create a new Python implementation from scratch it doesn't need to be there. It is not intrinsic to Python; it is intrinsic to the specific CPython implementation. It is also, broadly speaking, easy to remove the GIL. However the bar the CPython developers have set is that they will not accept a GIL-removal patch that harms single-threaded performance. This is the bar that has proved difficult to hurdle. It is also fair to point out that while it's a perennial topic in the Python community, it is not the only scripting language implementation with a GIL or moral equivalent in it. [1]: For similar reasons, none of the core implementations of the 1990s-era dynamic scripting languages have a very good true multi-threading story. (Some of the alternates can do it, like Jython.) They don't all even have implementations, and last I knew, none of them have anything that you ought to use in production. I don't think this is a fundamental limitation of any of the languages in question, it's just that it's really hard to retrofit threading onto a non-threaded code base after ~ten very heavy years of development. There's a set of other things that are hard to retrofit in to an existing code base if they don't start there from day one, like "unit testability" or "proper string handling".
- eloff 10y agoThis is not true - it is intrinsic to Python - not just CPython. Part of it is the garbage collector, reference counting is expensive to do without the GIL - but that can be changed. IronPython, Jython, some translations of PyPy all attest to that. The other part is C methods, which includes all operations on Python collections like lists and dicts, are atomic, making those collections safe to share and modify across threads. Most python code in the wild that utilizes threads makes use of those guarantees. So you either use fine-grained locking, or lock-free data structures, or STM, etc to protect ALL collection instances, or you've basically made a Python 3000 that breaks compatibility with all existing Python code. So every Python implementation that removes the GIL chooses to go the first route of pessimistically making everything threadsafe - which was tried in Java way back when and was and is a miserable failure. You give up too much performance for it to be worthwhile. So the problem in Python will never be fixed in my opinion - the problem is Python and there's no saving it. Without breaking backwards compatibility it will always be better to use processes instead of threads and keep the GIL.
- daveguy 10y agoI'm just glad that Guido hasn't gone on record saying there will never be an official python without a GIL.
- fatbird 10y agoGuido's on record as having three "political" requirements for any GIL-less Python: 1. Same or better performance 2. No breaking existing extensions 3. Not overly-complicating the cpython implementation These are tough requirements, but obviously sensible, and Hasting's discussion of trying to meet them is interesting.
- chadr 10y agoSemi related to this, is it possible to run multiple CPython interpreters in the same process, but with restrictions on shared memory? The idea being that each interpreter would still have its own GIL, each interpreter would have restrictions on shared memory (like sharing immutable structures only through message passing). Note, I'm not a big Python user so if this already exists, has been discussed, etc I am not aware.
- geofft 10y agoI'm not deeply familiar with the CPython API, but that should be completely straightforward as long as you don't trade native CPython structures between interpreters. If you start with something like Numpy's array interface, and implement an array in C that knows that it could be accessed by multiple Python interpreters and needs to become immutable before it's passed, it should work. But I wouldn't expect it to be easy to take a regular PyObject and move it between interpreters, so the cost of serialization and deserialization might be an issue.
- jamesdutc 10y agoSort of. Note that Python has support for shared memory: https://docs.python.org/2/library/multiprocessing.html#sharing-state-between-processes https://docs.python.org/2/library/multiprocessing.html#shari... In fact, `numpy` has its own mechanisms to support shared memory between processes: https://bitbucket.org/cleemesser/numpy-sharedmem https://bitbucket.org/cleemesser/numpy-sharedmem Neither of these approaches seem to be used very commonly in practice. Python itself has some "sub-interpreter" support. There was a long conversation about this last year: https://mail.python.org/pipermail/python-ideas/2015-June/034177.html https://mail.python.org/pipermail/python-ideas/2015-June/034... Finally, I have a working approach using `dlmopen` to host multiple interpreters within the same process: https://gist.github.com/dutc/eba9b2f7980f400f6287 https://gist.github.com/dutc/eba9b2f7980f400f6287 - the approach is so bizarre, because it's a very naïve multiple-embedding. It was intended to prove that you could run a Python 2 and a Python 3 together in the same process as part of a dare. This was thought impossible, since there are symbols with non-unique names that the dynamic linker would be unable to distinguish (which lead me to the `RTLD_DEEPBIND` flag for `dlopen`,) and that there is global state in a Python interpreter that interacts in undesirable ways (which lead me to `dlmopen` and linker namespaces.) - this approach is stronger than the traditional subinterpreter approach, since I can host multiple interpreters of distinct versions. i.e., I can host a Python 1.5 inside a Python 2.7 inside a Python 3.5. - the approach is stronger in that I completely isolate C libraries. There's a good amount of functionality provided by C libraries that maintain global state. e.g., `locale.setlocale` is a wrapping of C stdlib locale and is globally scoped. - this approach is weaker in that it requires a dynamic linker that supports linker namespaces, which effectively limits its use on Windows - this approach is weaker in that it's not complete: there's insufficient interest in this approach for me to actually write the shims to allow communication between processes. - this approach is weaker in that it has some weird restrictions such as being able to spawn only 15 sub-interpreters before running out of thread-local storage space I suppose the premise is that the GIL-removal efforts involve pessimistic coördination. A sub-interpreter approach might have a lighter touch and allow the user to handle coördination between processes (perhaps even requiring/allowing them to handle locks themselves.)
- jamesdutc 10y agoI had a personal conversation with Larry Hastings (the presenter) at PyCon. Here are a couple of notes from that chat, phrased neutrally. Some of these points may be reïterated in the linked video: - We can view this work is as a revisiting of Greg Stein's GIL-removal attempt in Python 1.4: http://dabeaz.blogspot.com/2011/08/inside-look-at-gil-removal-patch-of.html http://dabeaz.blogspot.com/2011/08/inside-look-at-gil-remova... It seems wholly reasonable to revisit the approach in light of how the language and ecosystem have changed since 1999. There are demands made of CPython core developers to remove or address the problem of the GIL, and these efforts demonstrate how much work is necessary to do that successfully. - Comparing single-threaded performance in a GIL implementation against single-threaded performance in a GIL-less implementation is considered an unfair comparison. A GIL-less will do extra book-keeping that will necessarily result in slower single-threaded performance.
- IanCal 10y ago> - Comparing single-threaded performance in a GIL implementation against single-threaded performance in a GIL-less implementation is considered an unfair comparison. A GIL-less will do extra book-keeping that will necessarily result in slower single-threaded performance. Unless you can choose between having the GIL or not, I think it's perfectly reasonable to compare the performance. If you can choose, then I think it's still useful to know the kind of overhead you're adding.
- pizlonator 10y agoI have some thoughts: - Doing better than atomic inc/dec for those reference counts is going to be hard. All of those techniques are still in an "unconfirmed myth" state AFAICT: someone published a paper but nobody has confirmed that the result holds on a broader set of machines, workloads, baselines, etc. - Great call on userspace locking. Note that you can do this portably. You don't need special OS support. See https://webkit.org/blog/6161/locking-in-webkit/ https://webkit.org/blog/6161/locking-in-webkit/ - Seems like lots of the locking use cases can indeed be made lock-free if you are willing to roll up your sleeves and get dirty. That's what I would do. - I still bet that the dominant cost is lock contention and he is not analyzing his data correctly. He appears to claim that it can't be locks because the total CPU time is greater than the total length of critical sections and that some analysis tools tell him that there is massive cache trashing. But that's exactly what happens if you contend on locks too often. Lock contention causes context switches and thread migrations. Both of those things require cache flushes. So the code that runs under contention will report massive cache thrashing because it will have a high probability of being on a cold cache. Programs indeed will run slower under contention than without it, and while his slow-down is extreme, I've seen worse and fixed it by removing some contention. He should find every contended lock and kill it with fire. - The dip at 4 cores doesn't surprise me. Computers are strange and most "scalability" charts (X axis is CPUs, Y axis is some measure of perf) I've made had weirdly reproducible dips and jerks.
- denfromufa 10y agoI asked about lock-free data structures from 2 major open-source libs, but Larry found major issues with them. He may consider some of these data structures in the future. I opened a bounty for this. See closed github issues for gilectomy.
- pizlonator 10y agoNot lock-free data structures. That's usually a fool's errand. You need lock-free hacks. Exhibit #1: the lock to protect lazy initialization. That just needs a CAS on the slow path and a load-load fence on the fast path. Delete the object you created if you lose the race and try again. Exhibit #2: you can probably do dirty things to make loading from dicts/lists not require locks even though storing to them does.
- hueving 10y agoWhat a nice talk. It's a shame the only question we got to hear him answer from the audience (only time left for 1) was such a lame troll. If you attend conferences, please don't waste everyone's time with nonsense. This is a python conference talking about a long-known perf issue and the guy just asked why not use another language. It's like going to WWDC and asking why everyone isn't using Android.
- airless_bar 10y agoI think it's a very valid question. Given that the constraints by the Python BDFL rule out any reasonable solution, why not move on?
- sametmax 10y agoThe problem is that the question is in the wrong context. Somebody is trying very hard to do something despite those constraints. For free. Why would you even try to discourage him to do so ?
- airless_bar 10y agoTo save him time and frustration due to his project going nowhere?
- sametmax 10y agoSince when trolling convinced somebody of anything? This was definitly not the intent of the troll, just a joke. But a badly executed one. It's like going to a political meeting, hearing somebody's plan to improve economy, and as a question asking him why not just change country. It makes little sense. Plus, everybody should consider the resume of the guy, cause he is not just a newcomer trying to show off. He is a hell of a good dev.
- airless_bar 10y agoI don't believe that the question was meant to be a joke or a trolling attempt. The guy's competence isn't the issue. It's the requirements of what the fix can and cannot do. _Everything_ is a mess due to the massive technical debt of never improving key parts of the runtime. I mean what's the point of a runtime where adding threads makes things magnitudes slower than running in a single-thread? "But that's just the state right now, they will fix this!" No, they won't. Just as a perspective: People usually fight over a single percent of improvement or less when working on runtime concurrency. Nobody just goes in and fixes _magnitudes_ of performance issues without rethinking what has been done and picking a completely different approach. This is not about making things "faster". This is literally going from doing necessary work to discovering a way of not having to do the work at all anymore. That's just not going to happen. Neither locks nor reference counting have anywhere near this optimization potential given the current language semantics. I think the whole (C)Python community missed the train more than a decade ago. Things that necessarily had to happen just didn't happen. Many communities and groups of developers "professionalize" over time. This doesn't seem to have happened in Python. Just an example: C extensions. It has been clear for at least a decade that the existing C interface won't work for threading, "real" garbage collection, etc. What would have been smart: Providing a better interface 10 years ago, effectively giving C ext devs 10 years of time to migrate their code. This way efforts to remove the GIL today would have to satisfy one fewer constraint that's currently crippling all of the efforts. Same with a lot of other things ... getting rid of the GIL would have required changes to various parts of the stack over the years (GC, language -> Python 3?, APIs ...), turning the actual removal of the GIL just into a final act of a multi-step process. What actually happened: Nothing. And now they just try to break the large lock into millions of smaller locks ... in 2016. WAT? This just tells me that key people in the Python community never really gave a damn about the issue and therefore this guy will waste his time. I have never expected Python's demise, but it now seems that largely its culture and not its lacking technology brought it down for good. From my perspective, they should have never released Python 3 with a GIL. That's what broke Python's neck, finally.
- andreasvc 10y agoIt seems to me that the option of going with a tracing garbage collector is preferable. Removing the GIL will require a lot of changes anyway, and it may be better to go all they way so as to preserve performance. It would affect the C API and extensions modules would have to be reworked for this, but on the other hand you avoid all the issues with reference counting. If extension modules rely on code generation, such as Cython, a move to a radically different C API might be less painful than expected. From what I know, most current managed languages are using GC, not reference counting, so perhaps it's an inherently better approach.
- Animats 10y agoMost of the Python implementations other than CPython use GC rather than reference counting. PyPy does. Iron Python did.[1] There is a version of PyPy without a GIL[2], but it runs much slower on ordinary code and is still under development. The developers are looking for financial support.[3] The approach is to identify large blocks of code as transactions, and run them in parallel. If they try to access the same data, one transaction fails and is backed out. It's like database rollback. But you have to write your code like this: from transaction import TransactionQueue tr = TransactionQueue() for key, value in bigdict.items(): tr.add(func, key, value) tr.run() [1] http://doc.pypy.org/en/latest/cpython_differences.html http://doc.pypy.org/en/latest/cpython_differences.html [2] http://doc.pypy.org/en/latest/stm.html http://doc.pypy.org/en/latest/stm.html [3] http://pypy.org/tmdonate2.html http://pypy.org/tmdonate2.html
- IanCal 10y agoFor the performance improvements, I think I'm missing something but why does coalesced reference counting have a high overhead? Is it normally implemented along with buffered reference counting? It feels like those fit together very neatly, one thread managing the counts and receiving updates from the other threads, and each other thread tries to only send updates it needs to. Is it simply a case of doing something basic a lot of times is faster than doing something smarter a few times because computers are just really fast at basic things? Or is there something more to this?
- ensiferum 10y agoLinux's pthread locks are already implemented on top of futex and have a fast user space path for the non-contended case.