9 ms·
The future of Python web services looks GIL-free
- mkoubaa 11mo agoI thought WSGI already used subinterpreters that each have their own GIL
- deleted 11mo ago[deleted]
- theandrewbailey 11mo agoSubinterpreters are part of the Python standard library as of 3.13 (I think?).
- ZiiS 11mo agoThis needs a lot of RAM with the speed/cores of modern CPUs. A GIL free multi threaded ASGI scales much further.
- pas 11mo agoWSGI is just the protocol for web request handling (like CGI, fast CGI), some implementations utilized subinterpreter support through the C API (which existed in its basic form since Python 1.5 according to PEP 554 or since 2.2 according to the current docs) but before 3.12 the isolation was not great (and still there are basic process-level things that cannot be non-shared per interpreter) https://docs.python.org/3/library/concurrent.interpreters.html#introduction https://docs.python.org/3/library/concurrent.interpreters.ht...
- Waterluvian 11mo agoWe already have PyPy and PyPI, so I think we are cosmically required to call Python 3.14 PyPi
- NeutralForest 11mo agoNice that someone takes the time to crunch the number, thanks! I know there's some community effort in how to use free-threaded Python: https://py-free-threading.github.io/ https://py-free-threading.github.io/ I've found debugging Python quite easy in general, I hope the experience will be great in free-threaded mode as well.
- shdh 11mo agoHadn’t heard of Granian before, thinking about upgrading to 3.14 for my services and running them threaded now
- Zsfe510asG 11mo agoAccessing a global object as the most simple benchmark that in fact exercises the locks still shows a massive slowdown that is not offset by the moderate general speedups since 3.9: x = 0 def f(): global x for i in range(100000000): x += i f() print(x) Results: 3.9: 7.1s 3.11: 5.9s 3.14: 6.5s 3.14-nogil: 8.4s That is a NOGIL slowdown of 18% compared to 3.9, 44% compared to 3.11 and 30% compared to 3.14. These numbers are in line with previous attempts at GIL removal that were rejected because they didn't come from Facebook. Please do not complain about the global object. Using a pure function would obviously be a useless benchmark for locking and real world Python code bases have far more intricate access patterns.
- logicchains 11mo ago>Please do not complain about the global object. Using a pure function would obviously be a useless benchmark for locking and real world Python code bases have far more intricate access patterns. Just because there's a lot of shit Python code out there, doesn't mean people who want to write clean, performant Python code should suffer for it.
- deleted 11mo ago[deleted]
- nine_k 11mo agoThose who's going to suffer are the people who inherited a ton of legacy Python code written in this style.
- NeutralForest 11mo agoAs mentioned in the article, others might have different constraints that make the GIL worth it for them; since both versions of Python are available anyways, it's a win in my book.
- mynewaccount00 11mo agoyou realize this is not a concern because nobody in the last 20 years uses global?
- throwaway984393 11mo ago[dead]
- natdempk 11mo agoReally great, just waiting on library support / builds for free threading. Have people had any/good experiences running Granian in prod?
- alex_hirner 11mo agoIt was and is a life saver. Our django app suffered from runaway memory leaks (quite a story). We were not able to track down the root cause exactly. There are numerous, similar issues with uvicorn or other webservers. Granian contained these problems. Multi process management is also reliable.
- pdhborges 11mo agoWhat did you try to debug this?
- alex_hirner 11mo agomemray and later a custom request wrapper that output python gc statistics. Our main candidates for the leak are: grpc, asgi server itself, psycopg, django channels. All leaked to some degree. Alas, it did not became clear what caused the runaway leak at 30 MB/s. Capturing flamegraphs just before OOM kills would require some more engineering. Granian contained these situations until upgradings later that year made the system more stable to begin with.
- sigwinch 11mo agoA board tracking progress on libraries: https://hugovk.github.io/free-threaded-wheels/ https://hugovk.github.io/free-threaded-wheels/
- Spivak 11mo ago> On asynchronous protocols like ASGI, despite the fact the concurrency model doesn't change that much – we shift from one event loop per process, to one event loop per thread – just the fact we no longer need to scale memory allocations just to use more CPU is a massive improvement. It's nice that someone else recognizes that event loop per thread is the way. I swear if you said this online any time in the past few years people looked at you like you insulted their mother. It's so much easier to manage even before the performance improvements.
- yupyupyups 11mo agoNo, because you can't kill a Python thread, but you can kill a process. That is a significant limitation to think about, especially when executing something long-running and CPU intensive such as large scientific computations. If your thread gets stuck, the only recourse you will have is to kill the entire parent process.
- tommmlij 11mo agoThat plus you can have memory leaks when you run heavy stuff in threads...
- hunterpayne 11mo agoWow, I seriously question the quality of any project where this is a consideration. I would also strongly recommend you hire some better devs and rewrite that project in another language with better concurrency features. Java (or anything on the JVM) would lead that list but plenty of languages would be suitable. Also, sharing memory between processes is very very very slow compared to sharing memory between threads.
- yupyupyups 11mo agoThe assumption is that they never share memory. Any time you have an arbitrary independent task that can be started and then stopped by the user, you will need processes in Python. Java is a lot better at concurrency and has a vast library for concurrency primitives like atomic operations (CAS etc.) Python's strengths lies in its strong ecosystem and ease of use. Many times that overshadows the benefits of Java's competent concurrency features.
- rogerbinns 11mo agoC code needs to be updated to be safe in a GIL free execution environment. It is a lot of work! The pervasive problem is that mutable data structures (lists, dict etc) could change at any arbitrary point while the C code is working with them, and the reference count for others could drop to zero if *anyone* is using a borrowed reference (common for performance in CPython APIs). Previously the GIL protected where those changes could happen. In simple cases it is adding a critical section, but often there multiple data structures in play. As an example these are the changes that had to be done to the standard library json module: https://github.com/python/cpython/pull/119438/files#diff-efe183ae0b85e5b8d9bbbc588452dd4de80b39fd5c5174ee499ba554217a39ed https://github.com/python/cpython/pull/119438/files#diff-efe... This is how much of the standard library has been audited: https://github.com/python/cpython/issues/116738 https://github.com/python/cpython/issues/116738 The json changes above are in Python 3.15, not the just released 3.14. The consequences of the C changes not being made are crashes and corruption if unexpected mutation or object freeing happens. Web services are exposed to adversity so be *very* careful. It would be a big help if CPython released a tool that could at least scan a C code base to detect free threaded issues, and ideally verify it is correct.
- westurner 11mo ago> It would be a big help if CPython released a tool that could at least scan a C code base to detect free threaded issues, and ideally verify it is correct. Create or extend a list of answers to: What heuristics predict that code will fail in CPython's nogil "free threaded" mode?
- rogerbinns 11mo agoSome of that is already around, but scattered across multiple locations. For example there is a list in the Python doc: https://docs.python.org/3/howto/free-threading-extensions.html#borrowed-references https://docs.python.org/3/howto/free-threading-extensions.ht... And a dedicated web site: https://py-free-threading.github.io/ https://py-free-threading.github.io/ But as an example neither include PySequence_Fast which is in the json.c changes I pointed to. The folks doing the auditing of stdlib do have an idea of what they are looking for, and so would be best suited to keep a list (and tool) up to date with what is needed.
- btbuilder 11mo agoThis is fantastic progress for CPython. I had almost given up hope that CPython would overcome the GIL after first hitting its limitations over 10 years ago. That being said I strongly believe that because of the sharp edges on async style code vs proper co-routine-based user threads like go-routines and Java virtual threads Python is still far behind optimal parallelism patterns.
- rowanG077 11mo agoAren't go-routines the worst of all worlds? Sharp edges, undefined behavior galore? At least that was my takeaway when last using about 5 or 6 years ago. Did they fix go-routines in the meantime?
- deleted 11mo ago[deleted]
- weakfish 11mo agoI like them, but I’m doing pretty simple back/end dev in Go, just microservices.
- btbuilder 11mo agoIt’s hard to answer without specifics but languages shouldn’t require you to determine whether it’s safe to use an api in an async context or whether it will hang your app. I imagine some of the sharp edges you might have run into are because go has real parallelism and you have to address data sharing.
- int_19h 11mo agoThe sharp edges in Go are when you try to use the built-in mechanisms that are supposed to replace data sharing - e.g. channels - only to discover the numerous footguns that abound there. And then there's patently stupid design decisions like using raw slices as collections and the maybe-change-maybe-copy semantics of append() that don't make it easier to reason about shared data when it needs to be shared.
- deleted 11mo ago
- callamdelaney 11mo agoPython gets more bloated weekly in my view.
- nine_k 11mo agoRemoving something, e.g. removing GIL, is usually the opposite of bloat.
- ViscountPenguin 11mo agoRemoving the GIL really amounts to adding a bunch of concurrency code all over th cPython codebase. It kind of sucks tbh.
- nodesocket 11mo agoI'm running a Python 3.13 Flask app in production using gunicorn and gevent (workers=1) with gevent monkey patching. Using ab can get around 320 requests per second. Performance is decent but I'm wondering how much a lift would be required to migrate to FastAPI. Would I see performance increases staying with gunicorn + gevent but upgrading Python to 3.14?
- stackskipton 11mo agoWe have similar at work, 3.14 should just be a Dockerfile change away. FastAPI might improve your performance by a little but seriously, either PyPy or rewriting into compiled language.
- nodesocket 11mo agoI did some quick tests increasing workers=2 and workers=3 and requests per second nearly scaled linearly so seems just throwing more CPU cores is the quick answer in the mid-term.
- nine_k 11mo agoDid you profile your code? Is it CPU-bound or IO-bound? Does it max out your CPU? Usually it's the DB access that determines the single-threaded performance of backend code.
- waldrews 11mo agoFor math/data-sci/ML, multiple (GIL-bound) interpreters per process, with (unsafely) shared data structures, would get us much of the way there - basically multiprocessing without the marshalling and process overhead, at the price of pinky-swearing we won't mutate the shared data. That would enable, for example, calling Python-world utilities in-process from a properly multi-threaded language, which can be used to bootstrap e.g. the Julia or Golang (or C#, or Rust...) ecosystems with all the math that's now locked into Python world. If we want NumPy/Scikit etc. to be accessible via thin wrappers, it's tolerable if the Python layer is slow, but importing the GIL into the host language is too high a price.