17 ms·
Free-threaded CPython is ready to experiment with
- nas 2y agoVery encouraging news!
- OutOfHere 2y agoIt has been ready for a few months now, at least since 3.13.0 beta 1 which released on 2024-05-08, although alpha versions had it working too. I don't know why this is news now. With it, the single-threaded case is slower.
- TylerE 2y agoFTA: "Yesterday, py-free-threading.github.io launched! It's both a resource with documentation around adding support for free-threaded Python, and a status tracker for the rollout across open source projects in the Python ecosystem."
- OutOfHere 2y agoBefore the article came the misleading title: "Free-threaded CPython is ready to experiment with". The link should have been to https://py-free-threading.github.io/tracking/ https://py-free-threading.github.io/tracking/
- JBorrow 2y agoThis release coincides with the SciPy 2024 conference and a number of other things. I would suggest reading the article to learn more.
- OutOfHere 2y ago> This release What release. The last release of CPython was 3.13.0b3 on 2024-06-27. SciPy is irrelevant to the title.
- nine_k 2y agoPython 3 progress so far: [x] Async. [x] Optional static typing. [x] Threading. [ ] JIT. [ ] Efficient dependency management.
- janice1999 2y agoNot sure what this list means, there are successful languages without these feature. Also Python 3.13 [1] has an optional JIT [2], disabled by default. [1] https://docs.python.org/3.13/whatsnew/3.13.html https://docs.python.org/3.13/whatsnew/3.13.html [2] https://peps.python.org/pep-0744/ https://peps.python.org/pep-0744/
- jolux 2y agoThe successful languages without efficient dependency management are painful to manage dependencies in, though. I think Python should be shooting for a better package management user experience than C++.
- yosefk 2y agoIf Python's dependency management is better than anything, it's better than C++'s. Python has pip and venv. C++ has nothing (you could say less than nothing since you also have ample opportunity for inconsistent build due to mismatching #defines as well as using the wrong binaries for your .h files and nothing remotely like type-safe linkage to mitigate human error. It also has an infinite number of build systems where each system of makefiles or cmakefiles is its own build system with its own conventions and features). In fact python is the best dependency management system for C++ code when you can get binaries build from C++ via pip install...
- wiseowise 2y ago> If Python's dependency management is better than anything, it's better than C++'s. That’s like the lowest possible bar to clear.
- infinghxsg 2y ago[dead]
- eigenvalue 2y agoReally excited for this. Once some more time goes by and the most important python libraries update to support no GIL, there is just a tremendous amount of performance that can be automatically unlocked with almost no incremental effort for so many organizations and projects. It's also a good opportunity for new and more actively maintained projects to take market share from older and more established libraries if the older libraries don't take making these changes seriously and finish them in a timely manner. It's going to be amazing to saturate all the cores on a big machine using simple threads instead of dealing with the massive overhead and complexity and bugs of using something like multiprocessing.
- phkahler 2y agoI feel like most things that will benefit from moving to multiple cores for performance should probably not be written in Python. OTH "most" is not "all" so it's gonna be awesome for some.
- MBCook 2y agoBut it would give you more headroom before rewriting for performance would make sense right? That alone could be beneficial to a lot of people.
- rty32 2y agoI think it is beneficial to some people, but not a lot. My guess is that most Python users (from beginners to advanced users, including many professional data scientists) have never heard of GIL or thought of doing any parallelization in Python. Code that needs performance and would benefit from multithreading, usually written by professional software engineers, likely isn't written in Python in the first place. It would make sense for projects that can benefit from disabling GIL without a ton of changes. Remember it is not trivial to update single threaded code to use multithreading correctly. in Python language specifically. Their library may have already done some form of parallelization under the hood
- 2y ago
- mihaic 2y agoDoes anyone know if there is more serious single threaded performance degradation (more than a few percent for instance)? I couldn't find any benchmarks, just some generic reassurance that everything is fine.
- ngoldbaum 2y agoRight now there is a significant single-threaded performance cost. Somewhere from 30-50%. Part of what my colleague Ken Jin and others are working on is getting back some of that lost performance by applying some optimizations. Expect single-threaded performance to improve for Python 3.14 next year.
- arp242 2y agoTo be honest, that seems a lot. Even today a lot of code is single-threaded, and this performance hit will also affect a lot of code running in parallel today. There have been patches to remove the GIL going back to the 90s and Python 1.5 or thereabouts. But the performance impact has always been the show-stopper.
- ngoldbaum 2y agoIt’s an experimental release in 3.13. Another example: objects that will have deffered reference counts in 3.14 are made immortal in 3.13 to avoid scaling issues from reference count thrashing. This wasn’t originally the plan but deferred reference counting didn’t land in time for 3.13. It will be several years before free-threading becomes the default, at that point there will no longer be any single-threaded performance drop. Of course that assumes everything shakes out as planned, we’ll see. This post is a call to ask people to “kick the tires”, experiment, and report issues they run into, not announcing that all work is done.
- andmkl 2y agoThat would be in the order of previous GIL-removal projects, which were abandoned for that reason.
- 2y ago
- Sparkyte 2y agoMy body is ready. I love python because the ease of writing and logic. Hopefully the more complicated free-threaded approach is comprehensive enough to write it like we traditionally write python. Not saying it is or isn't I just haven't dived enough into python multithreading because it is hard to put those demons back once you pull them out.
- ZhongXina 2y agoPrecisely, ease of writing, not ease of reading (the whole project, not just a tiny snippet of code) or supporting it long-term.
- ameliaquining 2y agoThe semantic changes are negligible for authors of Python code. All the complexity falls on the maintainers of the CPython interpreter and on authors of native extension modules.
- stavros 2y agoWell, I'm not looking forward to the day when I upgrade my Python and suddenly I have to debug a ton of fun race conditions.
- teaearlgraycold 2y agoIt's kept behind a flag. Hopefully will be forever.
- vldmrs 2y agoGreat news ! It would be interesting to see performance comparison for IO-bound tasks like http requests between single-threaded asyncio code and multi-threaded asyncio
- discreteevent 2y agoI remember back around 2007 all the anxious blog posts about the free lunch (Moore's law) being over. Parallelism was mandatory now. We were going to need exotic solutions like software transactional memory to get out of the crisis (and we could certainly forget about object orientation). Meanwhile what takes the crown? - Single threaded python. (Well, ok Rust looks like it's taking first place where you really need the speed and it does help parallelism without requiring absolute purity)
- jeremycarter 2y agoTakes what crown? Python is horrifically slow even single threaded. It's by far the slowest and most energy inefficient of the major choices available today.
- elijahbenizzy 2y agoI'm really curious to see how this will work with async. There's a natural barrier (I/O versus CPU-bound code), which isn't always a perfect distinction. I'd love to see a more fluid model between the two -- E.G. if I'm doing a "gather" on CPU-bound coroutines, I'm curious if there's something that can be smart enough to JIT between async and multithreaded implementations. "Oh, the first few tasks were entirely CPU-bound? Cool, let's launch another thread. Oh, the first few threads were I/O-bound? Cool, let's use in-thread coroutines". Probably not feasible for a myriad of reasons, but even a more fluid programming model could be really cool (similar interfaces with a quick swap between?).
- bastawhiz 2y agoI think you'd be hard pressed to find a workload where that behavior needs to be generalized to the degree you're talking. If you're serving HTTP requests, for instance, simply serving each request on its own thread with its own event loop should be sufficient at scale. Multiple requests each with CPU-bound tasks will still saturate the CPUs. Very little code teeters between CPU-bound and io-bound while also serving few enough requests that you have cores to spare to effectively parallelize all the CPU-bound work. If that's the case, why do you need the runtime to do this for you? A simple profile would show what's holding up the event loop. But still, the runtime can't naively parallelize coroutines. Coroutines are expected not to be run in parallel and that code isn't expected to be thread safe. Instead of a gather on futures, your code would have been using a thread pool executor in the first place if you'd gone out of your way to ensure your CPU-bound code was thread safe: the benefits of async/await are mostly lost. I also don't think an event loop can be shared between two running threads: if you were to parallelize coroutines, those coroutines' spawned coroutines could run in parallel. If you used an async library that isn't thread safe because it expects only one coroutine is executing at a time, you could run into serious bugs.
- elijahbenizzy 2y agoInteresting. I don't disagree, in general, but I actually have worked with a lot of applications that like to do this. Specifically in the world of ML/AI inference there's a lot of moving between external querying of data (features) and internal/external querying of models. With recommendation systems it is often worse -- gather large data, run a computation on it, filter it, get a bulk API request, score it with a model, etc... This is exactly where I'd like to see it. I'd like to simultaneously: 1. Call out to external APIs and not run any overhead/complexity of creating/managing threads 2. Call out to a model on a CPU and not have it block the event loop (I want it to launch a new thread and have that be similar to me) 3. Call out to a model on a GPU, ditto And use the observed resource CPU/GPU usage to scale up nicely with an external horizontal scaling system. So it might be that the async API is a lot easier to use/ergonomic then threads. I'd be happy to handle thread-safety (say, annotating routines), but as you pointed out, there are underlying framework assumptions that make this complicated. The solution we always used is to separate out the CPU-bound components from the IO-bound components, even onto different servers or sidecar processes (which, effectively, turn CPU-bound into IO-bound operations). But if they could co-exist happily, I'd be very excited. Especially if they could use a similar API as async does.
- anacrolix 2y agoWas ready for this 15 years ago when I loved Python and regularly contributed. At the time, nobody wanted to do it and I got bored and went to Go.
- throwaway5752 2y agoGVR, you are sorely missed, though I hope you are enjoying life.
- gnatolf 2y agoGood to hear. The authors are touching on the journey it is to make Cython continue to work. I wonder how hard it'll be to continue to provide bdist packages, or within what timeframe, if at all, Cython can transparently ensure correctness for a no-gil build. Anyone got any insights?
- jmward01 2y agoI know, I know, 'not every story needs to be about ML' but.... I can only imagine how unlocking the GIL will change the nature of ML training and inference. There is so much waste and complexity in passing memory around and coordinating processes. I know that libraries have made it (somewhat) easier and more efficient but I can't wait to see what can be done with things like pytorch when optimized for this.
- ipsum2 2y agoIt'll mostly help for debugging and lowering RAM (not VRAM) usage. Otherwise it won't impact ML much.
- jmward01 2y agoPretty universally I have seen performance improvements in code when complexity is reduced and this could drop complexity considerably. I wouldn't be surprised to see a double digit percent improvement in tokens per sec when an optimized pytorch eventually comes out with this. There may even be hidden gains on GPU memory usage that come out of this as people clean up code and start implementing better tricks because of it.
- imtringued 2y agoYeah, one of the dumbest things about Dataloaders running in a different process is that you are logging into the void.
- veber-alex 2y agohuh? Any python library that cares about performance is written in C/C++/Rust/Fortran and only provides a python interface. ML will have 0 benefit from this.
- jmward01 2y agoHave you done any multi-gpu training? Generally every GPU gets a process. Coordinating between them and passing around data between them is complex and can easily have performance issues since normal communication between python processes requires some sort of serialization/de-serialization of objects (there are many * here when it comes to GPU training). This has the potential to simplify all of that and remove a lot of inter-process communication which is just pure overhead.
- farhanhubble 2y agoIt remains to be seen how many subtle bugs are now introduced by programmers who have never dealt with real multithreading.
- simonw 2y agoI got this working on macOS and wrote up some notes on the installation process and a short script I wrote to demonstrate how it differs from non-free-threaded Python: https://til.simonwillison.net/python/trying-free-threaded-python https://til.simonwillison.net/python/trying-free-threaded-py...
- vanous 2y agoThanks for the example and explanations Simon!
- andmkl 2y ago[flagged]
- NegativeK 2y agoI downvoted you because your commend felt like it was a string of strawman arguments.
- kortex 2y agoMany uses. There's tons of situations where you are already accelerating most of your heavy compute with tensor libraries, but the data input/output parts are still in python. They would benefit from loading data in parallel before batching. Multiprocess, OS threads, and asyncio all solve different problems. Threads are pretty heavyweight compared to async coroutines (aka green threads). The big win with coroutines is it is very cheap to put them to sleep waiting on io. So a web server on a 4 core vm might have 4 worker processes, several threads per process, and dozens/hundreds of coroutines. > So perhaps you can use this for slurping other people's IP in parallel and train the "AIs" that are supposed to make us redundant. This has absolutely nothing to do with the technical merits of async or threads. For one, the above two examples are taken directly from work I did to combat deep fakes by identifying various "tells" from the media. Some of us are in fact using machine learning for good.
- pansa2 2y agoPEP703 explains that with the GIL removed, operations on lists such as `append` remain thread-safe because of the addition of per-list locks. What about simple operations like incrementing an integer? IIRC this is currently thread-safe because the GIL guarantees each bytecode instruction is executed atomically.
- pansa2 2y agoAh, `i += 1` isn’t currently thread-safe because Python does (LOAD, +=, STORE) as 3 separate bytecode instructions. I guess the only things that are a single instruction are some modifications to mutable objects, and those are already heavyweight enough that it’s OK to add a per-object lock.
- jillesvangurp 2y agoThat sounds like the kind of thing that a JIT compiler should be optimizing. The problem with threading isn't stuff like this but people doing a lot of silly things like having global mutable state or stateful objects that are being passed around a lot. I've done quite a bit of stuff with Java and Kotlin in the past quarter century and it's interesting to see how much things have evolved. Early on there were a lot of people doing silly things with threads and overusing the, at the time, not so great language features for that. But a lot of that stuff replaced by better primitives and libraries. If you look at Kotlin these days, there's very little of that silliness going on. It has no synchronized keyword. Or a volatile keyword, like Java has. But it does have co-routines and co-routine scopes. And some of those scopes may be backed by thread pools (or virtual thread pools on recent JVMs). Now that python has async, it's probably a good idea to start thinking about some way to add structured concurrency similar to that on top of that. So, you have async stuff and some of that async stuff might happen on different threads. It's a good mental model for dealing with concurrency and parallelism. There's no need to repeat two decades of mistakes that happened in the Java world; you can fast forward to the good stuff without doing that.
- vegabook 2y agoClearly the Python 2 to 3 war was so traumatising (and so badly handled) that the core Python team is too scared to do the obvious thing, and call this Python 4. This is a big fundamental and (in many cases breaking) change, even if it's "optional".
- blumomo 2y agoDid Python as the language change which justified that version bump?
- mixmastamyk 2y agoWhen on, there are incompatibilities yes. There were a lot of smaller breaking changes over the years, especially 3.10 that probably should have been a 4.0.
- grandimam 2y agoHow is the no-gil performance compared to other languages like - javascript (nodejs), go, rust, and even java? If it's bearable then I believe there is enormous value that could be generated instead of spending time porting to other languages.
- pansa2 2y agoNo-GIL Python is still interpreted - single-threaded performance is slower that standard Python, which is in turn much slower than the languages you mentioned. Maybe if you’ve got an embarrassingly parallel problem, and dozen(s) of cores to spare, you can match the performance of a single-threaded JIT/AOT compiled program.
- vulnbludog 2y agoHow do companies like Instagram/OpenAI scale with a majority python codebase? Like I just kick it on HN idk much about computers or coding (think high school CS) why wouldn’t they migrate can someone explain like I’m five
- pansa2 2y agoPython may well have been the right choice for companies like that when they were starting out, but now they're much bigger, they would be better off with a different language. However, they simply have too much code to rewrite it all in another language. Hence the attempts recently to fundamentally change Python itself to make it more suitable for large-scale codebases. <rant>And IMO less suitable for writing small scripts, which is what the majority of Python programmers are actually doing.</rant>
- imtringued 2y agoThey have tools like Triton that compile a restricted subset to CUDA.
- thebigspacefuck 2y agoHere’s a benchmark https://github.com/lip234/python_313_benchmark https://github.com/lip234/python_313_benchmark It’s much worse except in everything but a threaded test
- earthnail 2y agoOh how much this would simplify torch.DataLoader (and its equivalents)… Really excited about this.
- VagabundoP 2y agoHighly recommend the core.py podcast if you're interested in the background, there are a few episodes that focus on the GILectomy: -Episode 2: Removing the GIL[1] -Episode 12: A Legit Episode[2] [1]https://www.youtube.com/watch?v=jHOtyx3PSJQ&list=PLShJCpYUN3C3XrdguJiAZrQEnQGH-Sq26&index=13 https://www.youtube.com/watch?v=jHOtyx3PSJQ&list=PLShJCpYUN3... [2]https://www.youtube.com/watch?v=IGYxMsHw9iw&list=PLShJCpYUN3C3XrdguJiAZrQEnQGH-Sq26&index=1 https://www.youtube.com/watch?v=IGYxMsHw9iw&list=PLShJCpYUN3...
- codethief 2y agoYesterday someone presented preliminary benchmarks here at EuroPython 2024, comparing no-GIL to sub-interpreters and to multiprocessing. Upshot: This gon' be good!
- westurner 2y agoWill there be an effort to encourage devs to add support for free-threaded Python like for Python 3 [1] and for Wheels [2]? Is there a cibuildwheel / CI check for free-threaded Python support? Is there already a reason not to have Platform compatibility tags for free-threaded cpython support? https://packaging.python.org/en/latest/specifications/platform-compatibility-tags/ https://packaging.python.org/en/latest/specifications/platfo... Is there a hame - a hashtaggable name - for this feature to help devs find resources to help add support? Can an LLM almost port in support for free-threading in Python, and how should we expect the tests to be insufficient? "Porting Extension Modules to Support Free-Threading" https://py-free-threading.github.io/porting/ https://py-free-threading.github.io/porting/ [1] "Python 3 "Wall of Shame" Becomes "Wall of Superpowers" Today" https://news.ycombinator.com/item?id=4907755 https://news.ycombinator.com/item?id=4907755 [2] https://pythonwheels.com/ https://pythonwheels.com/ (Edit) Compatibility status tracking: https://py-free-threading.github.io/tracking/ https://py-free-threading.github.io/tracking/
- westurner 2y agoInstall commands from https://py-free-threading.github.io/installing_cpython/ https://py-free-threading.github.io/installing_cpython/ : sudo dnf install python3.13-freethreading sudo add-apt-repository ppa:deadsnakes sudo apt-get update sudo apt-get install python3.13-nogil conda create -n nogil -c defaults -c ad-testing/label/py313_nogil python=3.13 mamba create -n nogil -c defaults -c ad-testing/label/py313_nogil python=3.13 TODO: conda-forge ?, pixi
- westurner 2y ago(2021) https://news.ycombinator.com/item?id=29005573#29009072 https://news.ycombinator.com/item?id=29005573#29009072 : python-feedstock / recipe / meta.yml: https://github.com/conda-forge/python-feedstock/blob/master/recipe/meta.yaml https://github.com/conda-forge/python-feedstock/blob/master/... pypy-meta-feedstock can be installed in the same env as python-feedstock; https://github.com/conda-forge/pypy-meta-feedstock/blob/main/recipe/meta.yaml https://github.com/conda-forge/pypy-meta-feedstock/blob/main...