27 ms·
A viable solution for Python concurrency
- Animats 5y agoOr you could just use PyPy, which uses a garbage collector, does more compile-time analysis, and runs much faster. CPython is a naive interpreter, like original JavaScript. There's been progress since then.
- nerdponx 5y agoPyPy being stuck on 3.7 hurts. If 3.8 support comes out soon, I'll be happy to switch for general-purpose work. 3.9 would be even nicer, to support the type annotation improvements. I donate every month, but I'm just an individual donating pocket change; it'd be great to see some corporate support for PyPy.
- calpaterson 5y agoThere are very few new features in 3.8. It is a much less important release (for features) than 3.7, which for example added dataclasses and lots of typing and asyncio stuff. The most significant change in 3.8 is a notoriously controversial new infix operator. Even it's supporters would say that it's a niche usecase.
- masklinn 5y ago> There are very few new features in 3.8. > It is a much less important release (for features) than 3.7, which for example added dataclasses and lots of typing and asyncio stuff. That's funny because my take is the exact opposite: dataclasses are not very useful (attrs exists and does more), deferred type annotations are meh, contextvars, breakpoint(), and module-level getattr/settattr but not exactly anything you can't do without. Assignment expressions provide for great cleanups in some contexts (and avoiding redundant evaluations in e.g. comprehensions), expr= is tremendous for printf-debugging, posonly args is really useful, \N in regex can much improve their readability when relevant. $dayjob has migrated to python 3.7 and there's really nothing I'm excited to use (possibly aside from doing weird things with breakpoint), whereas 3.8 would be a genuine improvement to my day-to-day enjoyment.
- nerdponx 5y agoDeferred type annotations with `from __future__ import annotations` are a game-changer IMO. You can use them 3.7, which is good enough for me. The big improvement in 3.9 is not having to use `typing.*` for a lot of basic data types. The biggest improvements between 3.7, 3.8, 3.9, and 3.10 are in `asyncio`, which was pretty rough in 3.7 and very usable in 3.9. I use the 3rd-party `anyio` library in a lot of cases anyway (https://anyio.readthedocs.io/ https://anyio.readthedocs.io/), but it's not always feasible.
- llimllib 5y agolots of people need C extensions, which you can't* have on pypy. *: mostly true
- willvarfar 5y agoPypy is still single threaded. https://doc.pypy.org/en/latest/faq.html#does-pypy-have-a-gil-why https://doc.pypy.org/en/latest/faq.html#does-pypy-have-a-gil... This work is super exciting! Can pypy use the same recipe to offer true parallelism plus the jit?? Will be really interesting to see what pypy devs think of this work and how they might also lever it!
- nas 5y agoI think it can't use the same recipe. Sam's approach for CPython uses biased reference counting. Internally, Pypy uses a tracing garbage collector, not reference counting. I don't know how difficult it would be to make their GC thread-safe. Probably you don't want to "stop the world" on every GC pass so I guess changes are non-trivial. Sam's changes to CPython's container objects (dicts, lists), to make them thread safe might also be hard to port directly to Pypy. Pypy implements those objects differently.
- willvarfar 5y agoI think the biggest thing it will give is a need to go there. Until now, pypy has been able to not do parallelism. But if cpython is suddenly faster for a big class of program, pypy will have to bite the bullet to stay relevant?
- masklinn 5y agopypy also has a GIL.
- laurencerowe 5y agoIt's been a few years since I last played around with PyPy but while it provided amazing performance gains for simple algorithmic code I saw no speed up on a more complex web application.
- dehrmann 5y agoThere's also Jython. If it weren't for Python's love of Cython and Jython being stuck at 2.7, it would be very viable.
- overgard 5y agoI feel like Gvr just doesnt want to change things, Feels doomed This has been a problem for like 20 years and they have refused fixes before. And there have been fixes. They just don't see this as important it's practically a religion that its a thing they wont change
- heavyset_go 5y agoI disagree entirely. The last few releases of Python have made significant changes to the language, coinciding with the project becoming community-led after Guido stepped down.
- dragonwriter 5y ago> The last few releases of Python have made significant changes to the language, coinciding with the project becoming community-led after Guido stepped down. A lot of that is stuff that is enabled by, or was blocked pending, the new parser; I don't think it was blocked on Guido, and Guido haa hardly stopped being active and influential since stepping down as BDfL
- heavyset_go 5y agoYeah, I didn't mean to imply that Guido was holding anything back, just that he's no longer at the helm like the OP implied.
- mixmastamyk 5y agoYes, stepped down as a result of him forcing the walrus operator change into the language over significant opposition.
- randlet 5y agoGuido is no longer the BDF and spoke fairly positively about this change in the mailing list thread[1]. "To be clear, Sam’s basic approach is a bit slower for single-threaded code, and he admits that. But to sweeten the pot he has also applied a bunch of unrelated speedups that make it faster in general, so that overall it’s always a win. But presumably we could upstream the latter easily, separately from the GIL-freeing part." [1] https://mail.python.org/archives/list/python-dev@python.org/thread/ABR2L6BENNA6UPSPKV474HCS4LWT26GY/ https://mail.python.org/archives/list/python-dev@python.org/...
- the__alchemist 5y agoMaybe we just accept that Python isn't suitable for concurrency. There's a community of Python developers who don't want to branch out; write everything in Python and never learn or consider another language. Let them be. Let Python excel at its core competencies; use the right tool for the job.
- mixmastamyk 5y agoOdd time to write this comment as the first viable solution drops.
- rfoo 5y ago> Let Python excel at its core competencies Python is notably very popular in two communities: web developer and scientific computing. The former usually yell loudly every time someone propose to remove GIL. Meanwhile everyone in the scientific computing community had to learn how to workaround GIL which absolutely sucks and sometimes just impossible. (e.g. I have a mostly memory-bandwidth-bound data loading pipeline but it sometimes need multiple cores for doing some trivial data transformation with numpy. This is impossible to do efficiently in Python right now, due to GIL) Do you mean Python is not the right tool for my job and we should drop numpy/scipy/... and let Python be your shit tool for rendering web pages?
- detaro 5y agoWhy would webdevs care about work trying to remove the GIL?
- rfoo 5y agoPrevious GIL removal attempts hurts single thread performance and it isn't that scalable, so people are usually by default dismissive. Most of Python codes depend on subtle details of CPython internal. For example sometimes it is just convenient to assume GIL exists (i.e. simplifies concurrency codes because "you know there are at most one thread running").
- 5y ago
- misnome 5y ago> This "optimization" actually slows single-threaded accesses down slightly, according to the design document, but that penalty becomes worthwhile once multi-threaded execution becomes possible. My understanding was that CPython viewed any single-threaded performance regression as a blocker to GIL-removal attempts, regardless of if other work by the developer has sped up the interpreter? This article seems to somewhat gloss over that with "it's only small". I'd be interested in knowing other estimations of what the "better-than-average chance" of this (promising sounding) attempt were. Breaking C extensions (especially the less-conforming ones, which seem likely to be the least maintained) also seems like it would be a very hard pill to swallow, and the sort of thing that might make it a Python 3-to-4 breaking change, which I imagine would also be approached extremely carefully given there are still people to-this-day who believe that python 3 is a mistake and one day everyone will realise it and go back to python 2 (yes, really).
- fatbird 5y agoIt was Guido's requirement that GIL removal not degrade single threaded performance at all, but in the talk I attend at PyCon 2019, the speaker mentioned nothing about qualifications on that. Guido's restriction was presented, quite reasonably, as "no one should have to suffer because of removing the GIL". So a net break-even or performance improvement is fine. And on top of that, Guido has retired now, and the steering committee may feel differently as long as the spirit of the restrictions is upheld.
- fatbird 5y agoGuido has replied to Gross's announcement to observe that his performance improvements are not tied to removing the GIL and could be accepted separately. But he doesn't reject Gross's work outright, and if the same release that includes the GIL removal also delivers a concrete performance upgrade, I suspect that Guido would be fine with it. His concern is, after all, practical, to do with the actual use of python and not some architecture principle.
- cormacrelf 5y ago
- efoto 5y ago"The biggest source of problems might be multi-threaded programs with concurrency-related bugs that have been masked by the GIL until now."
- Animats 5y agoYes. I once discovered that CPickle was not thread-safe. The response was that much of the library didn't really work in multi-threaded programs.
- formerly_proven 5y agoYou mean programs where you put an object into pickle and some other threads modify it while pickle is processing it? Doesn't surprise me - the equivalent written in plain Python would be very thread unsafe as well.
- Animats 5y agoNo, I mean several threads doing completely separate CPickle streams with no shared data or variables at the Python level.
- kzrdude 5y agoHas it since been fixed?
- toyg 5y agoProbably not. CPickle is famously shunned by anyone who has to do serious, performance-critical serialization/deserialization.
- kzrdude 5y agoI was curious, and an issue that fits the description was fixed in Py 3.7.x here: https://bugs.python.org/issue34572 https://bugs.python.org/issue34572 but other threading bugs remain: https://bugs.python.org/issue38884 https://bugs.python.org/issue38884
- lormayna 5y agoWhy not using something trio or curio? They are quite easy to learn, very powerful and have an approach similar to channel in golang.
- synchronizing 5y agoasync != multithreading
- calpaterson 5y agoThose are not multithreading, they are asynchronous io which is different. With asynchronous io in Python the only concurrency/parallelism you can do is for IO. Multithreading in Python currently has the same limitation though but it needn't.
- rich_sasha 5y agoThis is some of the best news I read in a while! Multiprocessing sort of works but it’s really sucky.
- lucb1e 5y agoI actually have a great experience using that rather than dealing with concurrency hell. My use case is typically brute forcing this or that, for example a quick implementation to crack a key on some ctf challenge, or a proof of concept to crack a session token for a customer demo. Just spawn a few processes, each gets 1/nth of the work, not a big deal. But I could see how, if you want to have (e.g.) sound and UI rendering in completely independent "threads" (thus needing multiprocess) it could be a pain to link it all up. What kind of use case are you thinking of or did you run into where it was a really sucky experience?
- rich_sasha 5y agoFor me the main thing is, I cannot parallelise easily a loop mid-function. I need to make a pool, separate out the loop body into a separate top-level function, and also deal with multiprocessing quirks (like processes dying semi-randomly). It feels quite heavy and clunky, is my issue.
- kzrdude 5y agoEfficient threading is always like that anyway? You have to push the parallization a bit up in level. By the way, concurrent.futures in Python provides identical (almost) features for a process and thread pool executor, so either choice there can be handled the same way. I'm most fond of the executor.map() method for very easy parallelization of work-loops.
- singhrac 5y agoNotably the dev proposing this (Sam Gross aka colesbury) is/was a major PyTorch developer, so someone quite familiar with high performance Python and extensions.
- ajtulloch 5y agoand he's a genius!
- jeremyis 5y ago+1 ! Though, not a very good Oculus player (yet)!
- alfanerd 5y ago+1
- benreesman 5y agoComing from you that’s very high praise legend! (Tulloch is the best hacker and mathematician I’ve worked with in a 20 year career).
- ferdowsi 5y agoIf this effort succeeds (and I hope it does) now Python developers will need to contend with the event-loop albatross of asyncio and all of its weird complexity. In an alternate Python timeline, asyncio was not introduced into the Python standard library, and instead we got a natively supported, robust, easy-to-use concurrency paradigm built around green/virtual threading that accommodates both IO and CPU bound work.
- harpiaharpyja 5y agoIf you are ever considering making use of asyncio for your project, I would strongly recommend taking a look at curio [1] as an alternative. It's like asyncio but far, far easier to use. [1] https://curio.readthedocs.io/en/latest/index.html https://curio.readthedocs.io/en/latest/index.html
- acidbaseextract 5y agoThe video (or blog post) below is one of the best explanations I've seen about what subtle bugs are easy to make with asyncio, why it's easy to make them, and how the trio library addresses them. But yes, consider alternatives before you pick asyncio as your approach! Talk: https://www.youtube.com/watch?v=oLkfnc_UMcE https://www.youtube.com/watch?v=oLkfnc_UMcE Blog post: https://vorpus.org/blog/notes-on-structured-concurrency-or-go-statement-considered-harmful/ https://vorpus.org/blog/notes-on-structured-concurrency-or-g...
- VWWHFSfQ 5y agoHighly recommend curio
- deleted 5y ago[deleted]
- BiteCode_dev 5y agoWhile the design of Curio is quite interesting, it's may not be a good choice, not for technical reasons, but for logistical reasons: the chances it gets a wide adoption are slim to None. And since we are stuck with colored functions in python, the choice of stack matters very much. Now, if you want easier concurrency, and a solution to a lot of concurrency problems that curio solves, while still being compatible with asyncio, use anyio: https://anyio.readthedocs.io/en/stable/ https://anyio.readthedocs.io/en/stable/ It's a layer that works on top of asyncio, so it's compatible with all of it. But it features the nursery concept from Trio, which makes async programming so much simpler and safer.
- dsr_ 5y agoI'm going to assume that there is a reason that this isn't a switch control, so that the default is a single-threaded program and the programmer needs to state explicitly that this one will be multi-threaded, upon which the interpreter changes into the atomic mode for the rest of execution?
- toxik 5y agoThat would be expensive.
- mikepurvis 5y agoBasically no one would get the glorious single-threaded performance then, since the first time you pip install anything, you're going to discover that it spins up a thread under the hood that you're never exposed to. Or worse, you end up with the async schism all over again, with new "threadless" versions of popular libraries springing up.
- __s 5y agoMost references are thread local, where this implementation will still beat out atomic refcounts in a multi-threaded app
- mzs 5y agoYikes, C extensions can't assume they are under GIL by default: https://github.com/colesbury/numpy/commits/v1.19.3-nogil https://github.com/colesbury/numpy/commits/v1.19.3-nogil
- kzrdude 5y agoIt looks like a total of four lines needed changing in numpy due to his change. That's a very good score in my book, numpy is huge.
- Jweb_Guru 5y agoUnfortunately, every C extension will need to undergo manual review for safety, unless there's some very easy way to have the C extension opt into using the GIL. And some of them will be close to impossible to detangle in this way.
- veryupwork 5y agono
- int_19h 5y agoIt really depends on how the library is written, and how much shared data it has. It has been very common to use GIL as a general-purpose synchronization mechanism in native Python modules, since you have to pay that tax either way.
- phkahler 5y ago>> If that bit is set, the interpreter doesn't bother tracking references for the relevant object at all. That avoids contention (and cache-line bouncing) for the reference counts in these heavily-used objects. This "optimization" actually slows single-threaded accesses down slightly, according to the design document, but that penalty becomes worthwhile once multi-threaded execution becomes possible. Was going to say do the opposite. Set the bit if you want counting and then modify the increment and decrement to add or subtract the bit, thereby eliminating condition checking and branching. But it sounds like the concern is cache behavior when the count is written. Checking the bit can avoid any modification at all.
- Jtsummers 5y agoUnder this scheme, objects get freed when both local and shared counts are zero. By using a special value that makes the shared count non-zero (for eternal and long-lived objects), it ensures that should the owner (for some reason) drop them, they will not be freed. No extra logic has to be introduced, the shared count is non-zero and that's all that's needed to prevent freeing.
- nomdep 5y agoIf the Python maintainers doesn’t want to approve this, Gross should talk to the Pypi developers.
- dragonwriter 5y ago> If the Python maintainers doesn’t want to approve this, Gross should talk to the Pypi developers. I suspect you mean “PyPy” which is a very different thing than “pypi”.
- Ericson2314 5y ago> Gross has also put some significant work into improving the performance of the CPython interpreter in general. Earmarks work, folks!
- a1369209993 5y ago> With this scheme, the reference count in each object is split in two, with one "local" count for the owner (creator) of the object and a shared count for all other threads. Since the owner has exclusive access to its count, increments and decrements can be done with fast, non-atomic instructions. Any other thread accessing the object will use atomic operations on the shared reference count. > Whenever the owning thread drops a reference to an object, it checks both reference counts against zero. If both the local and the shared count are zero, the object can be freed, since no other references exist. If the local count is zero but the shared count is not, []a special bit is set to indicate that the owning thread has dropped the object[]; any subsequent decrements of the shared count will then free the object if that count goes to zero. This seems... off. Wouldn't it work better for the owning thread to hold (exactly) one atomic reference, which is released (using the same decref code as other threads) when the local reference count goes to zero? Edit: I probably should have explicitly noted that, as jetrink points out, the object is initialized with a atomic refcount of one (the "local refcount is nonzero" reference), and destroyed when the atomic refcount is one and to-be-decremented, so a purely local object never has atomic writes.
- cogman10 5y agoYeah, I'm not exactly getting all the complexity here. I'm digging the 2 reference counters, that makes sense to me, but I don't know why it isn't something more like: "every time a new thread takes a reference, atomic +1, every time a new thread's local count hits 0, atomic -1. If the shared reference is 0, free". IDK what special purpose the flags are serving here.
- deleted 5y ago[deleted]
- jeremyjh 5y agoMost objects are never shared so there would be a performance impact from incrementing (and decrementing) an atomic counter even just once.
- 5y ago
- jeremyis 5y agoWay to go Sam! Mark my words: our generations' Carmack!
- cormacrelf 5y ago> "biased reference counts" and is described in this paper by Jiho Choi et al. With this scheme, the reference count in each object is split in two, with one "local" count for the owner (creator) of the object and a shared count for all other threads > The interpreter's memory allocator has been replaced with mimalloc These are very similar ideas! Mimalloc is notable for its use of separate local and remote free lists, where objects that are being freed from a different thread than the page’s heap’s owner are placed in a separate queue. The local free list is (IIRC) non-atomic until it is empty and local allocs start pulling from the remote queue. The general idea is clearly lazy support for concurrency, matching up perfectly with Python’s need to keep any single threaded perf it has. I’m impressed with the application of all of these things at once.
- ikiris 5y agoI was half expecting a link to go or rust.
- The_rationalist 5y agoOr you could just use GraalVM python https://github.com/oracle/graalpython https://github.com/oracle/graalpython
- typical182 5y agoThis is a great list of influences on the design (from the article comments where the prototype author Sam Gross responded to someone wishing for more cross pollination across language communities): ————— "… but I'll give a few more examples specific to this project of ideas (or code) taken from other communities: - Biased reference counting (originally implemented for Swift) - mimalloc (originally developed for Koka and Lean) - The design of the internal locks is taken from WebKit (https://webkit.org/blog/6161/locking-in-webkit/ https://webkit.org/blog/6161/locking-in-webkit/) - The collection thread-safety adapts some code from FreeBSD (https://github.com/colesbury/nogil/blob/nogil/Python/qsbr.c https://github.com/colesbury/nogil/blob/nogil/Python/qsbr.c) - The interpreter took ideas from LuaJIT and V8's ignition interpreter (the register-accumulator model from ignition, fast function calls and other perf ideas from LuaJIT) - The stop-the-world implementation is influenced by Go's design (https://github.com/golang/go/blob/fad4a16fd43f6a72b6917eff656be27522809074/src/runtime/proc.go#L1154-L1230 https://github.com/golang/go/blob/fad4a16fd43f6a72b6917eff65... )"
- deleted 5y ago[deleted]
- marris 5y agoHow big a problem is the possible breakage of C extensions for new code? Is there currently some standard "future proofed for multi-thread" way of writing them that will reduce the odds of the C extension breaking? And maybe also being compatible with PyPy? Or do developers today need to write a separate version for each interpreter that they want to support?
- heavyset_go 5y agoThere are projects[1] that are abstracting away the C extension interface in order to standardize C extensions across implementations and prevent breaking changes. [1] https://github.com/hpyproject/hpy https://github.com/hpyproject/hpy
- sandGorgon 5y agohas anyone built and run this in docker ? would love to test this out - i dont have a lot of experience in compiling python inside docker EDIT: there is a dockerfile in there https://raw.githubusercontent.com/colesbury/nogil/nogil/Dockerfile https://raw.githubusercontent.com/colesbury/nogil/nogil/Dock...
- twic 5y ago> With this scheme, the reference count in each object is split in two, with one "local" count for the owner (creator) of the object and a shared count for all other threads. Since the owner has exclusive access to its count, increments and decrements can be done with fast, non-atomic instructions. Any other thread accessing the object will use atomic operations on the shared reference count. > Whenever the owning thread drops a reference to an object, it checks both reference counts against zero. If both the local and the shared count are zero, the object can be freed, since no other references exist. If the local count is zero but the shared count is not, a special bit is set to indicate that the owning thread has dropped the object; any subsequent decrements of the shared count will then free the object if that count goes to zero. So in this program: import threading def produce(): global global_foo local_foo = "potato" global_foo = local_foo def consume(): global global_foo local_foo = global_foo global_foo = None if __name__ == '__main__': produce() thread = threading.Thread(target=consume) thread.start() thread.join() What happens to the counts on the string "potato"? In produce, the main thread creates it and puts it in a local, and increments the local count. It assigns it to a global, and increments the local count. It then drops the local when produce returns, and decrements the local count. In consume, the second thread copies the global to a local, and increments the shared count. It clears out the global, and decrements the shared count. It then drops the local when consume returns, and decrements the shared count. That leaves the local count at 1 and the shared count at -1! You might think that there must be special handling around globals, but that doesn't fix it. Wrap the string in a perfectly ordinary list, and put the list in the global, and you have the same problem. I imagine this is explained in the paper by Choi et al, but i have not read it!
- Jtsummers 5y agoTwo spaces in front of each line of the code block. As written, right now, your comment is hard to parse: import threading def produce(): global global_foo local_foo = "potato" global_foo = local_foo def consume(): global global_foo local_foo = global_foo global_foo = None if __name__ == '__main__': produce() thread = threading.Thread(target=consume) thread.start() thread.join()
- otterley 5y agoThis may be a silly question, but if you really need concurrency, why not use a language that's built for concurrency from the ground up instead? Elixir is a great example.
- KZerda 5y agoLots of reasons. Sometimes concurrency is a relatively soft need where you could get by with multiple processes, but it would be nice if the language itself provided some capability for things like threads. Or your dev team is much more familiar with python than with other languages, and the time to retrain and rewrite would be greater than the benefits. Or the rest of the language could have greater issues and you don't want to give up excellent libraries like numpy.
- ska 5y agoTo a first approximation, people don't use python for itself, they use it for the vast ecosystem and network effect. If you jump to another language for better concurrency, what are you giving up? Unless you really are doing greenfield development in an isolated application, these considerations often trump any language feature.
- otterley 5y agoDon't get me wrong; I'm not suggesting that anyone dump Python altogether to switch to a different language for any arbitrary project or purpose. Many businesses I work with use different languages for different components or applications, using the network or storage (or even shared memory) to intercommunicate when necessary. The right tool for the job, as it were.
- deleted 5y ago[deleted]
- Mehdi2277 5y agoI use python mostly for numpy/tensorflow/etc. The machine learning ecosystem. That ecosystem cares a lot about multithreaded performance. So historically answer has been write c/c++ and then bind to python. This work is mainly motivated by libraries like that wanting to write less c extensions and be able to just write python and still have proper thread performance.
- Waterluvian 5y agoI can see it now: “My program has 2^64 references to an object, which caused it to become immortal” =)
- notriddle 5y agoIn a 64-bit address space, with objects requiring more than one word to store, that’s literally impossible.
- jeffybefffy519 5y agoI dont know about others, but I really enjoy content about the Python GIL. Its a fascinatingly complex problem.
- lucb1e 5y agoFor a minute I thought I finally found someone else who likes the GIL, but then you said content about. Programs that just divide up work across processes are much easier to write without introducing obscure bugs due to the lack of atomicity. I'm definitely excited for a GIL-less python, even if it's a rare scenario where it makes sense to try to do performant code in python in the first place rather than offloading a few lines to another language to be fast, but I am a bit afraid that people (particularly beginner programmers) will grab this with too many hands. Having also seen recommendations for this-or-that threading method going around in other languages, threads are recommended really much more often than where it makes sense and beginners won't have a comparative experience yet of writing multi-process code instead. That said, I am also always interested in GIL-related content like this! Loved the article.
- globular-toast 5y ago> Programs that just divide up work across processes are much easier to write without introducing obscure bugs due to the lack of atomicity. You often don't even need to do this yourself. GNU parallel is the way to go for dividing work up amongst CPU cores. Why reinvent the wheel? I agree with you that threads are talked about way more than they should be. It's like all programmers learn this one simple rule: to be fast you have to be multi-threaded. It's really not the case. There is also massive confusion amongst programmers on the difference between concurrency and parallelism. I sometimes ask applicants to describe the difference and few can. Python is fine at concurrency if that's all you want to do.
- adwn 5y ago> GNU parallel is the way to go for dividing work up amongst CPU cores. Why reinvent the wheel? Because most problems are not the embarrassingly parallel kind suitable for use with GNU parallel. For example, any problems that require some communication between the individual tasks.
- metalliqaz 5y agoEvery time multithreading and the GIL comes up, I wonder why there are so many out there that are against multiprocess. In addition to solving the GIL problem by simply having multiple GILs, it also forces the designer to think properly about inter-thread data flow. Sure, it can never be quite as fast as true multithreading, but the results are probably more robust and as a bonus it doesn't break all the existing Python libraries.
- kortex 5y agoMultiprocess isn't a panacea. Several frameworks get grumpy depending on the order of forking. E.g. Iirc, try to use grpc to feed data to a process using pytorch dataloader, and it'll straight up crash. Huggingface is at least polite enough to warn you, but performance degrades.
- pansa2 5y agoIs CPython the only widely-used language implementation that uses reference counting rather than tracing garbage collection? IIRC even PyPy and MicroPython use tracing GC.
- MichaelMoser123 5y agoperl5 uses reference counted garbage collection (since perl 5.8 - came out in 2002). They avoid the GIL by having a seperate interpreter instance per thread https://perl.mines-albi.fr/perl5.8.5/5.8.5/sun4-solaris/threads.html#description https://perl.mines-albi.fr/perl5.8.5/5.8.5/sun4-solaris/thre... If you want to share variables between threads, then they have to be marked as shared https://perl.mines-albi.fr/perl5.8.5/5.8.5/sun4-solaris/threads/shared.html https://perl.mines-albi.fr/perl5.8.5/5.8.5/sun4-solaris/thre...
- rzimmerman 5y agoObjective C does (explicit) reference counting, though I suppose that's a pretty limited use case these days.
- loeg 5y agoDoesn't Swift (built on ObjC) also use reference counting?
- sgerenser 5y agoYes, Swift uses refcounting.
- pansa2 5y agoDoes Swift have a GIL? If not, how does it solve the problem of multithreaded reference counting?
- loeg 5y agoPer earlier comment[1], Swift uses the biased refcounting approach Gross is proposing for Python. [1]: https://news.ycombinator.com/item?id=28882299 https://news.ycombinator.com/item?id=28882299
- huetius 5y agoMany of the additions to Python in the past decade have been very impressive, but am I wrong in thinking that they suffer from a kind of diminishing marginal benefit? If I am building a project where concurrency or asynchrony are essential, am I going to choose Python? If I need to bolt these on to an existing project to meet a deadline, how much runway do I really get from these enhancements before I hit the limitations of the language, and need a new tool anyway? I love Python, and use it a great deal, but I’ve not seen anything in my experience to assuage these doubts. Happy to be convinced otherwise, though. I’m sure there are circumstances I haven’t thought of.
- unityByFreedom 5y agoThis is more about making parallel python easier to use. It's already fast if you know about the current GIL workarounds. It'd be nice to not have to monkey patch, for example.
- huetius 5y agoQuite true. Definitely a benefit. I guess my point is that if you are leaning on Python for performant concurrent operations, you are likely to also be thinking about how to isolate that component for a rewrite, monkey-patching or no.
- unityByFreedom 5y ago> how to isolate that component At the most basic level you only need to put the logic into a function. Of course, the cost of that adjustment may vary, so more options are good. If Python becomes easier to use in this respect then everyone wins.
- pansa2 5y ago> am I going to choose Python? IMO most of the additions to Python in the past decade haven’t been aimed at people choosing a language, but at people working on large existing Python codebases who have too much inertia to change languages. The cause of this is that the people in charge of language design (first Guido at Dropbox and now members of the Steering Council) are in exactly that situation. The effect is that Python has lost focus on what it was originally good for (executable pseudocode for scripting) in favour of adding mediocre support for concurrency, static typing, and now pattern matching.
- sydthrowaway 5y agoHow does one get this good as a developer!
- hatware 5y agoI always imagine these are the folks who started programming as teenagers or earlier. Or, incredibly seasoned vets into their 50's and beyond.
- sydthrowaway 5y agoOr wealthy parents (feeds into 1)
- mrfusion 5y agoCould this justify a 4.0 versioning? 1. It’s a huge change and 2. There’s a small chance to break existing code where the gil was masking problems.
- mrfusion 5y agoI’d think there are a certain class of programs that could benefit from just turning off garbage collection and reference counting. If you’re sure you won’t be needing a lot of memory and the program will be short lived that could be a large speed up and you could skip the Gil.
- ttul 5y agoIs there a Nobel category for “major Python improvements”?
- george_ciobanu 5y agoFree Dask cluster available to use at tcp://3.216.44.221:8786 tldr: you can use our 4-computer Dask cluster (including two GPUs, a GTX 1080Ti and a Radeon 6900XT) at no cost. If you need even more computing power, message me (george AT lindenhoney DOT com). FAQ What is Dask? Dask is a Python framework for distributed computing, designed to enable data scientists to process large amounts of data and huge computations on a scale from 1 to thousands of computers. How do I use the cluster you created? Install Dask (python -m pip install "dask[complete]" from dask.distributed import Client client = Client(tcp://3.216.44.221:8786) Use it! https://examples.dask.org https://examples.dask.org Do I need to sign up? No. You can literally just use it as above from any Python script or notebook. Why are you making it available? I’m looking to better understand how people use Dask and how to best make it available to data scientists as an easy to use service. What can I use the cluster for? Anything you want such as data science, web crawling (except spam, porn, illegal activities or DoS) etc. Please be nice. Can I hack your cluster? Probably not. But ff you find any security issues please let us know instead and we will fix them. What’s the catch? Do you track my usage? None. These are my gaming/closet computers that I just want people to use. I do track anonymous function calls people make and the total volume of data that passes through the system (bandwidth, disk usage, RAM usage etc) but I do not look at or save any of your data. How can I add my own worker machines to the cluster? Easy - set up Dask worker(s) and point them to tcp://3.216.44.221:8786 How about security and confidentiality? These are computers I have at home so all you have is my promise that I won’t look/save/copy your data or code. Pinky swear. However, it is possible for other unfriendly users to look at your data if you save it, even temporarily, on the disk, since all Dask tasks run under the same user. Who are you? https://www.linkedin.com/in/georgeciobanunyc https://www.linkedin.com/in/georgeciobanunyc What if someone else who uses the cluster messes up with my data, copies it or otherwise takes it down? Use your best judgement. I hardened the cluster’s security to the best of my ability but I can’t guarantee that someone else won’t mess with them. If you work for the NSA/a bank/IRS/etc you should definitely not upload sensitive data to these computers. I need more GPUs or a bigger (hundreds or thousands of computers) cluster! Glad to help, just email me (see above). I need a larger cluster but private (using AWS, Azure, GCP) Same as above, email me and I can help. What are the computer specs? Two of them are AMD processors with 8 physical cores each, the other is a quad core Intel and the last one is a Mac (Intel, 4 cores). The dedicated graphics cards (GPUs) are on the AMD computers.
- game_the0ry 5y agoI might have a controversial or unpopular opinion - I do not think python should try to be concurrent or any more performant than it already is. Python is the tool I pick for quick scripting, not for highly performant systems - there are languages and run times for that. Sometime highly productive software does not need to be highly performant, and that does not make the software any more or less valuable. And sometimes speed of iteration is more valuable than speed of run time execution.
- ferdowsi 5y agoThat's a fair opinion. Like many academics and data scientists are probably fine with Python as a tool for scripting or as a simple interface language for calling C libraries. But if Python continues as is, it will continue to to lose major enterprise traction as web companies transition to using faster languages for applications and infrastructure. It'll be tough to stave off the negative feedback loop at that point.
- Twisol 5y agoI'm not confident that's a bad thing, and I worry that we're judging languages on metrics that only make sense for startups. Does a language need to chase growth? I'm not saying Python shouldn't improve. But, if it comes to it, Python shouldn't cannibalize the niche it filled to do so. That will just spark other languages to fill the niche again.
- pansa2 5y ago> Python shouldn't cannibalize the niche it filled It’s too late for that, IMO. Python’s already added async-await because it wants to be C#, type hints because it wants to be Java, and pattern matching because it wants to be Haskell. People who liked Python because of what made it different from other languages have long since been left behind.
- ledauphin 5y agoright or wrong, Python is hardly a niche language at this point. There's really nothing to cannibalize. Also, making threading better is not necessarily a new feature if done well - it could be mostly transparent to users.
- mmastrac 5y agoPython3 would have been a great time to _also_ break the C interface in a way that would make multi-threading easier. An opt-in for C libraries that are multi-threaded-aware could be useful as well. It would be a forcing function to ensure that libraries _eventually_ become MT-aware and eventually the older versions would drop away.
- SoylentOrange 5y agoPython 3.0 was released in 2008, over 13 years ago. We are almost certainly much closer to python 4.0 than to 3.0 today (given 3.10 RC is currently live)
- rmbyrro 5y agoCurrent version being 3.10 doesn't make it any closer to 4. It can go to 3.99. And they actually started talking about being able to go even further beyond that before a 4.0.
- codetrotter 5y agoRelevant: https://www.techrepublic.com/article/programming-languages-why-python-4-0-will-probably-never-arrive-according-to-its-creator/ https://www.techrepublic.com/article/programming-languages-w... > "I'm not thrilled about the idea of Python 4 and nobody in the core dev team really is – so probably there never will be a 4.0 and we'll just keep numbering until 3.33, at least," he said in a video Q&A. but also: > Van Rossum didn't rule out the possibility of Python 4.0 entirely, though suggested this would likely only happen in the event of major changes to compatibility with C. "I could imagine that at some point we are forced to abandon certain binary or API compatibility for C extensions… If there was a significant incompatibility with C extensions without changing the language itself and if we were to be able to get rid of the GIL [global interpreter lock]; if one or both of those events were to happen, we probably would be forced to call it 4.0 because of the compatibility issues at the C extension level," he said.
- SoylentOrange 5y agoThat’s true you’re right. But I’ve heard others talk about adding a JIT, which would likely be a large breaking change at the ABI level.
- cletus 5y agoThis feels like Schrodinger's Cake to me (you know, having it and eating it too). > ... the first of which is called "biased reference counts" ... With this scheme, the reference count in each object is split in two, with one "local" count for the owner (creator) of the object and a shared count for all other threads. Since the owner has exclusive access to its count, increments and decrements can be done with fast, non-atomic instructions. Any other thread accessing the object will use atomic operations on the shared reference count. So if the owner needs to check the shared ref count every time it changes the count of the non-atomic local count, isn't that basically all the negatives of a single shared atomic counter? Every design has pros and cons. Of course the GIL has well-known costs but it also has benefits. It makes developing C modules that integrate into Python relatively safe and trivial. And this is where Python shines: as plumbing where CPU intensive work (which is still single-threaded) is done by C modules. This is how the likes of Mercurial, Jupyter, numpy and scipy. And probably PyTorch but I know less about that. My personal view is that the world has largely moved on from dynamically typed languages for anything non-trivial or that isn't essentially plumbing. For good reason. Of course people will bring up Javascript but it has a captive market as being the only thing that'll universally run code in a browser and the likes of TypeScript can ease that burden anyway. This just feels fighting the seemingly inevitable fate of Python.
- DSingularity 5y agoThe shared ref count is in the same cache line making the check basically free. Once you shift to shared then it is equivalent to the atomic shared ref count.
- SoylentOrange 5y ago> My personal view is that the world has largely moved on from dynamically typed languages for anything non-trivial or that isn't essentially plumbing. For good reason. Of course people will bring up Javascript but it has a captive market as being the only thing that'll universally run code in a browser and the likes of TypeScript can ease that burden anyway. This is not correct. There is a lot of data analysis code in the research and scientific community written in python. A lot of PyCon attendees and speakers come from these communities. Oftentimes, it’s not easy to write code to perform a task entirely in numpy, and then you incur massive slowdowns (often 20-100x). This is a common and contemporary problem in the python ecosystem. GVR initially didn’t think that python needed to be faster either, but recently changed his mind. You can find a presentation on his motivations here: https://github.com/faster-cpython/ideas/blob/main/FasterCPythonDark.pdf https://github.com/faster-cpython/ideas/blob/main/FasterCPyt... Sam Gross comes out of that community so he’s familiar with the motivations for making raw python faster.
- chaostheory 5y agoPython’s multiprocessing library is more than good enough for me
- OOPMan 5y agoIt's weird how the more I work with Python, the less I want to work with Python. I moved into the language full-time in 2010 and it's now 2021. The packaging ecosystem is still a burning dumpster fire, the performance is still hot garbage and the whole approach to asyncio makes me want to bang my head against a wall. Tthe latest additions in Python 3.10 have me shaking my head. I love pattern matching (Yes, Scala fanboy detected) but shoving them into Python just seems....poorly thought out. I really hope to move away from the language in the long-term because I feel like it's a bad thing when I would rather work in Java or C++ than Python. For me it feels like Java and C++ took a look at themselves at the end of the 2000s and said "Okay, we need to sort something out, what we're doing now is not winning any hearts" while Python did also did some introspection and decided "Meh, let's just keep throwing mud at the wall until something sticks". It's one of the few languages I've worked with which seems to be actively getting worse every year, which is kinda sad :-/
- rmdashrfstar 5y agoYou’ll find camaraderie in the Rust community, we have similar stories… :) One of us, one of us, one of us!
- OOPMan 5y agoRust is on my list of things to look at one day, but I'm still on the Scala train for now ;-)
- garmaine 5y agoAnother former Python dev, now rust dev… give it a try. I have trouble putting it in words, but Rust has the feeling of ease of expressiveness that makes Python fun to work with, but with a top notch static typing system. Lots of former Python devs doing rust work now.
- OOPMan 5y agoSounds like similar reasons why I enjoy Scala. It feels like Python, but with a proper type system :-)
- xiaodai 5y agoWhy no mention of Julia yet?
- modeless 5y agoThreads are certainly important, but I have to say that I found the multiprocessing package to work very well. I think a lot of the things people think they need threads for would actually be better with multiprocessing instead. Memory protection is good! Shared memory is still available and explicit sharing of just what you need is better in a lot of ways than implicit sharing of everything. I will be glad if the GIL is fixed but I think people reach for threads too quickly and too often.
- azinman2 5y agoIf the GIL is fixed, why would you want multiprocessing versus threads? Threads are cheaper to create, easier to communicate between (even if you need to be careful), and simply do different things than what multiprocessing intends (eg easier for blocking I/O on many threads, versus multiprocessing which is really more of a task queue)
- modeless 5y agoWhy don't we just run all code in different threads of the same process? Multiple processes are more robust to failure and easier to reason about because they are less tightly coupled and more explicit about sharing. You can do blocking I/O with multiprocessing, you just have to explicitly share buffers.
- rfoo 5y ago> Multiple processes are more robust to failure and easier to reason about Yeah of course, except when you use multiprocessing to do some CPU-intensive work and one of the subprocess get OOMKilled. Now your main process just hang forever. This was reported on bpo roughly 10 years ago and the response is "we can't fix this". Is this enough to convince you multiple processes sometimes lead to unnecessary complexity?
- kabber 5y agoDon't fall for it Python users - Lucy will pull the football away just before you kick it.
- qstwall 5y agoIt is hard to see the point of a GIL removal that will destabilize the C extension ecosystem for probably a decade again: C extensions can already start as many threads as they like. Threaded pure Python, even if GIL-less, is still slow, so what is the point? The whole point of Python (before asyncio, pattern matching etc.) was being simple and having a nice C interface. If that continues to erode, people will (and should) look at other languages. C++ for example is pretty Pythonic these days, Java does not have these problems, Erlang (while slow) was written for concurrency from scratch, etc.
- profquail 5y agoNow could be a good time to make this change, in coordination with HPy: https://github.com/hpyproject/hpy https://github.com/hpyproject/hpy I agree though — it’s tempting to keep extending and stretching the language to be something it was never designed for; but at some point it’s been stretched so far it loses the properties that made it attractive to start with. I like Python, but some of the things people are using it for now, they should really consider another language instead, and write a Python wrapper on top of that if they must use it from Python.
- nmca 5y agoI work a non-FB $BIGCORP with lots of python super-experts who are very excited about this. Sam is deeply familiar with pytorch and C extensions; compatibility here seems to be the real deal. High hopes!
- rylact 5y agoI'm completely new to pytorch (I looked at it for several hours today). To put it in a neutral manner, it seems to have a rather high tolerance for complexity and relying on dozens of external packages, including pybind11. Memory leaks included, as a cursory Google search reveals. I hope this new style of writing Python packages does not leak into the interpreter.
- sandGorgon 5y agoI think this is a GREAT time to be doing this. This will undoubtedly shake the c-extension ecosystem. But it is already going to shake up because of Hpy (https://lwn.net/Articles/851202/ https://lwn.net/Articles/851202/) So might as well do it in one shot. But what im really interested in is - if this can be ported to Pypy. Given that pypy is already quite a bit faster than cpython...it would be interesting to see what the nogil will unleash
- MaXtreeM 5y agoSlight off topic but I am curious about using bits in an integer for flags. As the article mentions Gross uses 2 least significat bits for flags and the rest is an integer for reference counting. When someone considers whether to use most significant bits or least significant bits are there any major differences? Is is easier to implement or faster because of processor architectures/instruction sets to use least significnt bits or is that just a matter of choice?
- tzs 5y ago> Is is easier to implement or faster because of processor architectures/instruction sets to use least significnt bits or is that just a matter of choice? I have no idea if this is the reason in this particular in case, but on most architectures if you can fit everything you need to atomically modify together into one word there will be an instruction to do that quickly. For example many architectures have a compare-and-swap instruction. That generally takes 3 arguments: a memory address, an old value, and a new value. It atomically compares the word at the given address to the old value, and if and only if they are the same writes the new value to the memory address.
- ryanpetrich 5y agoA benefit of storing the reference count in the high bits is that overflows will never corrupt the flag bits and can be detected using the processor's flags instead of requiring a separate check. I'm not sure if this property is used here.
- deleted 5y ago[deleted]
- Tronchenbiais 5y agoRegarding the implementation of this, I am surprised by the use of the least significant bits of the local ref count to hold what are essentially flags telling whether the reference is immortal or deferred. This sounds like a ugly hack, and as pointed out by the article, will break existing C code manipulating the count directly without using the associated macros (although arguably, doing this is unspecified, and really it's their fault). My guess is that this was done for memory efficiency, but is there really no way to have the local refcount be a packed struct or something, where the "count" field holds the real ref count and the flags are stored separately? I have no understanding of the CPython internals, so it may very well be impossible, but I would appreciate someone explaining why.