9 ms·
Python 3.15's JIT is now back on track
- oystersareyum 7mo ago> We don’t have proper free-threading support yet, but we’re aiming for that in 3.15/3.16. The JIT is now back on track. I recently read an interview about implementing free-threading and getting modifications through the ecosystem to really enable it: https://alexalejandre.com/programming/interview-with-ngoldbaum/ https://alexalejandre.com/programming/interview-with-ngoldba... The guy said he hopes the free-threaded build'll be the only one in "3.16 or 3.17", I wonder if that should apply to the JIT too or how the JIT and interpreter interact.
- zarzavat 7mo agoI continue to believe that free-threading hurts performance more than it helps and Python should abandon it. Having to have thread safe code all over the place just for the 1% of users who need to have multi-threading in Python and can't use subinterpreters for some reason is nuts.
- kzrdude 7mo agoI don't want to go too heavy on the negatives, but what's nuts is Python going for trust-the-programmer style multithreading. The risk is that extension modules could cause a lot of crashes.
- gwking 7mo agoMy understanding is that many extension modules are already written to take advantage of multithreading by releasing the GIL when calling into C code. This allows true concurrency in the extension, and also invites all the hazards of multithreading. I wonder how many bugs will be uncovered in such extensions by the free threaded builds, but it seems like the “nuts” choice actually happened a long time ago.
- pansa2 7mo agoMaybe they could have two versions of the interpreter, one that’s thread-safe and one that’s optimised for single-threading? Microsoft used to do this for their C runtime library.
- veber-alex 7mo agoThat's exactly what we have now and it looks like the python devs want a single unified build at some point
- chuckadams 7mo agoPHP does this as well. Most distributions ship PHP without thread safety, but it's seeing more use now that FrankenPHP uses it. Speaking of which, it would be nice if PHP's JIT got a little love: it's never eked out more than marginal gains in heavily-numeric code.
- cpgxiii 7mo ago> Having to have thread safe code all over the place just for the 1% of users who need to have multi-threading in Python and can't use subinterpreters for some reason is nuts. Way more than 1% of the community, particularly of the community actively developing Python, wants free-threaded. The problem here is that the Python community consists of several different groups: 1. Basically pure Python code with no threading 2. Basically pure Python with appropriate thread safety 3. Basically pure Python code with already broken threaded code, just getting lucky for now 4. Mixed Python and C/C++/Rust code, with appropriate threading behavior in the C or C++ components 5. Mixed Python and C or C++ code, with C and C++ components depending on GIL behavior Group 1 gets a slightly reduced performance. Groups 2 and 4 get a major win with free-threaded Python, being able to use threading through their interfaces to C/C++/Rust components. Group 3 is already writing buggy code and will probably see worse consequences from their existing bugs. Group 5 will have to either avoid threading in their Python code or rewrite their C/C++ components. Right now, a big portion of the Python language developer base consists of Groups 2 and 4. Group 5 is basically perceived as holding Python-the-language and Python-the-implementations back.
- zarzavat 7mo agoWhere is the major win? Sorry but I just don't see the use case for free-threading. Native code can already be multi-threaded so if you are using Python to drive parallelized native code, there's no win there. If your Python code is the bottleneck, well then you could have subinterpreters with shared buffers and locks. If you really need to have shared objects, do you actually need to mutate them from multiple interpreters? If not, what about exploring language support for frozen objects or proxies? The only thing that free threading gives you is concurrent mutations to Python objects, which is like, whatever. In all my years of writing Python I have never once found myself thinking "I wish I could mutate the same object from two different threads".
- cpgxiii 7mo ago> Native code can already be multi-threaded so if you are using Python to drive parallelized native code, there's no win there. When using something like boost::python or pybind11 to expose your native API in Python, it is not uncommon to have situations where the native API is extensible via inheritance or callbacks (which are easy to represent in these binding tools). Today with the GIL you are effectively forced to choose between exposing the native API parallelism or exposing the native API extensibility; e.g. you can expose a method that performs parallel evaluation of some inputs, OR you can expose a user-provided callback to be run on the output of each evaluation, but you cannot evaluate those inputs and run a user-provided callback in parallel. The "dumbest" form of this is with logging; people want to redirect whatever logging the native code may perform through whatever they are using for logging in Python, and that essentially creates a Python callback on every native logging call that currently requires a GIL acquire/release. Could some of this be addressed with various Python-specific workarounds/tools? Probably. But doing so is probably also going to tie the native code much more tightly to problematic/weird Pythonisms (in many cases, the native library in question is an entirely standalone project). > The only thing that free threading gives you is concurrent mutations to Python objects, which is like, whatever. The big benefit is that you get concurrency without the overhead of multi-process. Shared memory is always going to be faster than having to serialize for inter-process communication (let alone that not all Python objects are easily serializable).
- zadikian 7mo agoPure Python code always needed mutexes for thread safety with or without ol' GIL. I thought the difficulty with removing the GIL instead had to do with C extensions that rely on it.
- superbatfish 7mo agoThis is accurate and the parent commenter here seems to be echoing a common misconception. Either they are confused or they need to elaborate more to demonstrate that they have a valid complaint. For instance, this would have been a valid complaint: "Users who don't need free threading will now suffer a performance penalty for their single-threaded code." That is true. But if you are currently using multiple threads, code that was correct before will still be correct in the free threaded build, and code that was incorrect before will still be incorrect.
- zadikian 7mo agoI got it wrong too, even after reading the docs, until I just tried it for myself. Intuitively, why would Python have threads that can't run fully in parallel but also create race conditions? Seems like the worst of both worlds, and it almost is, except you can still speed up io-bound or even GIL-releasing CPU-bound C calls this way. Some also mix it up with async JS which also can't use multiple CPUs and guarantees in-order execution until you "await" something. Well now that's asyncio in Python. Doesn't help that so much literature uses muddy terms like "async programming."
- reinhash 7mo agoI also wonder how many people actually need free-threading. And I wonder how useful it will be, when you can already use the ABI to call multi-threaded code. I think the GIL provides python with a great guarantee, I would probably prefer single-thread performance improvements over multithreading in python to be honest. Anyway if I need performance, Python would probably not be my first choice
- ekjhgkejhgk 7mo agoDoesn't PyPy already have a jit compiler? Why aren't we using that?
- olivia-banks 7mo agoAs far as I know, PyPy doesn't support all CPython extensions, so pure Python code will probably (very likely) run fine but for other things most bets are off. I believe PyPy also only supports up to 3.11?
- JoshTriplett 7mo agoBecause PyPy seems to be defunct. It hasn't updated for quite a while. See https://github.com/numpy/numpy/issues/30416 https://github.com/numpy/numpy/issues/30416 for example. It's not being updated for compatibility with new versions of Python.
- LtWorf 7mo ago[flagged]
- Waterluvian 7mo agoIt supports at best Python 3.11 code, right? So it’s not unmaintained, no. But the project is currently under resourced to keep up with the latest Python spec.
- LtWorf 7mo agoThat is not the same thing at all, and not what he said.
- JoshTriplett 7mo agoIt is exactly what I'm referring to. I didn't say there aren't still people around. But they're far enough behind CPython that folks like NumPy are dropping support. Unless they get a substantial injection of new people and new energy, they're likely to continue falling behind.
- rafph 7mo ago[flagged]
- rsoto2 7mo agoI am trying to push back. I don't care if other people think the tools make them faster, I did not sign up to be a guinea pig for my employer or their AI-corp partner.
- adrian17 7mo agoI'm been occasionally glancing at PR/issue tracker to keep up to date with things happening with the JIT, but I've never seen where the high level discussions were happening; the issues and PRs always jumped right to the gritty details. Is there anywhere a high-level introduction/example of how trace projection vs recording work and differ? Googling for the terms often returns CPython issue tracker as the first result, and repo's jit.md is relatively barebones and rarely updated :( Similarly, I don't entirely understand refcount elimination; I've seen the codegen difference, but since the codegen happens at build time, does this mean each opcode is possibly split into two (or more?) stencils, with and without removed increfs/decrefs? With so many opcodes and their specialized variants, how many stencils are there now?
- flakes 7mo agoYou’ll probably want to look to the PEPs. Havent dug into this topic myself but looks related https://peps.python.org/pep-0744/ https://peps.python.org/pep-0744/
- adrian17 7mo agoI think CPython already had tier2 and some tracing infrastructure when the copy-and-patch JIT backend was added; it's the "JIT frontend" that's more obscure to me.
- sheepscreek 7mo agoUPDATE: I misunderstood the question :-/ You can ignore this. I love playing with compilers for fun, so maybe I can shed some light. I’ll explain it in a simplified way for everyone’s benefit (going to ignore the stack): When an object is passed between functions in Python, it doesn’t get copied. Instead, a reference to the object’s memory address is sent. This reference acts as a pointer to the object’s data. Think of it like a sticky note with the object’s memory address written on it. Now, imagine throwing away one sticky note every time a function that used a reference returns. When an object has zero references, it can be freed from memory and reused. Ensuring the number of references, or the “reference count” is always accurate is therefore a big deal. It is often the source of memory leaks, but I wouldn’t attribute it to a speed up (only if it replaces GC, then yes).
- ecshafer 7mo agoWhat is wrong with the Python code base that makes this so much harder to implement than seemingly all other code bases? Ruby, PHP, JS. They all seemed to add JITs in significantly less time. A Python JIT has been asked for for like 2 decades at this point.
- stmw 7mo agoSome languages are much harder to compile well to machine code. Some big factors (for any languages) are things like: lack of static types and high "type uncertainty", other dynamic language features, established inefficient extension interfaces that have to be maintained, unusual threading models...
- RussianCow 7mo agoThat makes sense if you're comparing with Java or C#, but not Ruby, which is way more dynamic than Python. The more likely reason is that there simply hasn't been that big a push for it. Ruby was dog slow before the JIT and Rails was very popular, so there was a lot of demand and room for improvement. PHP was the primary language used by Facebook for a long time, and they had deep pockets. JS powers the web, so there's a huge incentive for companies like Google to make it faster. Python never really had that same level of investment, at least from a performance standpoint. To your point, though, the C API has made certain types of optimizations extremely difficult, as the PyPy team has figured out.
- vlovich123 7mo agoGoogle, Dropbox, and Microsoft from what I can recall all tried to make Python fast so I don’t buy the “hasn’t seen a huge amount of investment”. For a long time Guido was opposed to any changes and that ossified the ecosystem. But the main problem was actually that pypy was never adopted as “the JIT” mechanism. That would have made a huge difference a long time ago and made sure they evolved in lock step.
- int_19h 7mo ago
- fluidcruft 7mo ago(what are blueberry, ripley, jones and prometheus?)
- max-m 7mo agoThe names of the benchmark runners. https://doesjitgobrrr.com/about https://doesjitgobrrr.com/about
- fluidcruft 7mo agoSo the biggest gains so far are on Windows 11 Pro of (x86_64) ~20%? Is that because Windows was bad as a baseline (promethius)? It doesn't seem like the x86_64/Linux has improved as dramatically ~5% (ripley). I'm just surprised OS has that much of an effect that can be attributed to JIT vs other OS issues.
- raddan 7mo agoIt's hard to say whether it's Windows related since the two x86_64 machines don't just run different OSes, they also have different processors, from different manufacturers. I don't know whether an AMD Ryzen 5 3600X versus Intel i5-8400 have dramatically different features, but unlike a generic static binary for x86_64, a JIT could in principle exploit features specific to a given manufacturer.
- mkl 7mo agoYes, the graphs are incomprehensible because those are not defined in the article. They turn out to be different physical machines with different architectures: https://doesjitgobrrr.com/about https://doesjitgobrrr.com/about blueberry (aarch64) Description: Raspberry Pi 5, 8GB RAM, 256GB SSD OS: Debian GNU/Linux 12 (bookworm) Owner: Savannah Ostrowski ripley (x86_64) Description: Intel i5-8400 @ 2.80GHz, 8GB RAM, 500GB SSD OS: Ubuntu 24.04 Owner: Savannah Ostrowski jones (aarch64) Description: Apple M3 Pro, 18GB RAM, 512GB SSD OS: macOS Owner: Savannah Ostrowski prometheus (x86_64) Description: AMD Ryzen 5 3600X @ 3.80GHz, 16GB RAM OS: Windows 11 Pro Owner: Savannah Ostrowski
- 7mo ago
- killingtime74 7mo agoSorry but the graphs are completely unreadable. There are four code names for each of the lines. Which is jit and which is cpython?
- mkl 7mo agoThey are all JIT on different architectures, measured relative to CPython. https://doesjitgobrrr.com/about https://doesjitgobrrr.com/about: blueberry is aarch64 Raspberry Pi, ripley is x86_64 Intel, jones is aarch64 M3 Pro, prometheus is x86_64 AMD.
- killingtime74 7mo agoThanks
- deleted 7mo ago[deleted]
- AgentMarket 7mo ago[flagged]
- anon291 7mo agoReference counting is not a strict requirement for python. Certainly not accurate counting.
- 1819231267 7mo ago[flagged]
- jqbd 7mo agoWait is this real? Does it mean this person read it or the bot read it, I don't think this is moltbook if the latter
- ayhanfuat 7mo agoAgentMarket is a bot spamming multiple threads with AI generated comments, if that is what you are asking.
- AgentMarket 7mo ago[flagged]
- wei03288 7mo ago[flagged]
- owaislone 7mo agoOh man, Python 2 > 3 was such a massive shift. Took almost half a decade if not more and yet it mainly changing superficial syntax stuff. They should have allowed ABIs to break and get these internal things done. Probably came up with a new, tighter API for integrating with other lower level languages so going forward Python internals can be changed more freely without breaking everything.
- gjvc 7mo agoyes. it was not a massive shift. it was barely worth the effort.
- pansa2 7mo agoThe Python devs didn’t want to make huge changes because they were worried Python 3 would end up taking forever like Perl 6. Instead they went to the other extreme and broke everyone’s code for trivial reasons and minimal benefit, which meant no-one wanted to upgrade. Even the main driver for Python 3, the bytes-Unicode split, has unfortunately turned out to be sub-optimal. Python essentially bet on UTF-32 (with space-saving optimisations), while everyone else has chosen UTF-8.
- rjh29 7mo agoIronically Perl 5 managed to do the bytes-Unicode split with a feature gate, no giant major version change.
- diziet_sma 7mo ago> Python essentially bet on UTF-32 (with space-saving optimisations) How so? Python3 strings are unicode and all the encoding/decoding functions default to utf-8. In practice this means all the python I write is utf-8 compatible unicode and I don't ever have to think about it.
- sheept 7mo agoUTF-32 allows for constant time character accesses, which means that mystr[i] isn't O(n). Most other languages can only provide constant time access for code units.
- vanderZwan 7mo ago> However, I misunderstood and came up with an even more extreme version: instead of tracing versions of normal instructions, I had only one instruction responsible for tracing, and all instructions in the second table point to that. Yes I know this part is confusing, I’ll hopefully try to explain better one day. This turned out to be a really really good choice. I found that the initial dual table approach was so much slower due to a doubling of the size of the interpreter, causing huge compiled code bloat, and naturally a slowdown. > By using only a single instruction and two tables, we only increase the interpreter by a size of 1 instruction, and also keep the base interpreter ultra fast. I affectionally call this mechanism dual dispatch. I really do hope they'll write that better explanation one day because this sounds pretty intriguing all on its own.
- thunky 7mo agoI always wanted this for Python but now that machines write code instead of humans I feel like languages like Python will not be needed as much anymore. They're made for humans, not machines. If a machine is going to do the dirty work I want it to produce something lean, fast, and strictly verified.
- JodieBenitez 7mo agoPretty much my thoughts the other day... now that Codex does the writing, maybe I can finally switch to Go for the web backend stuff without being annoyed by some of its archaisms and gain significant execution performance, while still having a relatively easy to read language.
- kccqzy 7mo agoYou ask a machine to write your code and you still care about being easy to read? In my experience the people who care the most about code readability tend to be the people most opinionated on having the right abstractions, which are historically not available in Go.
- thunky 7mo agoI don't think people mind reading Go as much as they mind writing it.
- kccqzy 7mo agoNah all the `if err != nil` is just so much noise they obscures the real logic. And for the longest time it didn’t have generics to write map/filter/reduce on slices, forcing people to use loops where the intention is less clear.
- maleldil 7mo agoIdeally, the errors shouldn't be returned as-is, but wrapped with context instead. If that context doesn't matter for you, you can have your editor wrap the if instead, which helps a lot.
- ghm2199 7mo agoThanks for all the amazing work! I have Noob question. Wouldn't this get the funding back? Or would that not be preferable way to continue(as opposed to just volunteer driven)? Like this is a big deal to get a project to a state where volunteers are spun up and actively breaking tasks and getting work done, no? It's a python JIT something I know next to nothing about — as do most application developers — which tells one how difficult this must have been.
- pansa2 7mo ago> Wouldn't this get the funding back? The funding was Microsoft employing most of the team. They were laid off (or at least, moved onto different projects), apparently because they weren't working on AI.
- kelvinjps 7mo agoWith Python being the main language for AI, isn't like more important to be more performant? I kinda don't get Microsoft reasoning, maybe they're just tight in money
- brianwawok 7mo agoI don’t think Python is the main language of AI.
- eru 7mo agoPython is pretty big as glue in the AI ecosystem as far as I can tell. It also seems to be most agent's 'preferred' language to write code in, when you don't specify anything. (The latter is probably more to do with the preferences they give it in the re-inforcement learning phase than anything technical, though.)
- Ralfp 7mo agoIt looks like ARM picked up plenty of those folk and pays them to continue this work.
- rslashuser 7mo agoI'm curious is the JIT developers could mention any Python features that prevent promising JIT features. An earlier Ken Jin blog [1], mentions how __del__ complicates reference counting optimization. There is a story that Python is harder to optimize than, say, Typescript, with Python flexibility and the C API getting mentioned. Maybe, if the list of troublesome Python features was out there, programmers could know to avoid those features with the promise of activating the JIT when it can prove the feature is not in use. This could provide a way out of the current Python hard-to-JIT trap. It's just a gist of an idea, but certainly an interesting first step would be to hear from the JIT people which Python features they find troublesome. [1] https://fidget-spinner.github.io/posts/faster-jit-plan.html https://fidget-spinner.github.io/posts/faster-jit-plan.html
- rtpg 7mo agoIt's interesting you mention __del__ because Javascript not only doesn't have destructors but for security reasons (that are above my pay grade) but the spec _explicitly prohibits_ implementations from allowing visibility into garbage collection state, meaning that code cannot have any visibility into deallocations. I think __del__ is tricky though. In theory __del__ is not meant to be reliable. In practice CPython reliably calls it cuz it reference counts. So people know about it and use it (though I've only really seen it used for best effort cleanup checks) In a world where more people were using PyPy we could have pressure from that perspective to avoid leaning into it. And that would also generate more pressure to implement code that is performant in "any" system.
- nvme0n1p1 7mo ago> code cannot have any visibility into deallocations Doesn't FinalizationRegistry let you do exactly that? https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/FinalizationRegistry https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
- rtpg 7mo agoOh! While this one does mention that you don't have visibility, this + weak refs seem to change the game I remember a couple of years ago (well probably around 2021) reading about GC exposure concerns and seeing some line in some TC39 doc like "users should not have visibility into collection" but if we've shipped weakrefs sounds like we're not thinking about that anymore
- aplomb1026 7mo ago[dead]
- fivedicks 7mo ago[flagged]
- deleted 7mo ago[deleted]
- devnotes77 7mo ago[dead]
- openclaw01 7mo ago[dead]
- mattclarkdotnet 7mo agoPython really needs to take the Typescript approach of "all valid Python4 is valid Python3". And then add value types so we can have int64 etc. And allow object refs to be frozen after instantiation to avoid the indirection tax. Sensible type-annotated python code could be so much faster if it didn't have to assume everything could change at any time. Most things don't change, and if they do they change on startup (e.g. ORM bindings).
- panzi 7mo agoIsn't rpython doing that, allowing changes on startup and then it's basically statically typed? Does it still exist? Was it ever production ready? I only once read a paper about it decades ago.
- mattclarkdotnet 7mo agoRPython is great, but it changes semantics in all sorts of ways. No sets for example. WTF? The native Set type is one of the best features of Python. Tuples also get mangled in RPython.
- zahlman 7mo agoIt exists in the sense that PyPy exists. As far as I can tell, it only ever existed to make PyPy possible, and was only defined/specified in terms of PyPy's needs.
- mattclarkdotnet 7mo agoTo clarify, it is nuts that in an object method, there is a performance enhancement through caching a member value. class SomeClass def init(self) self.x = 0 def SomeMethod(self) q = self.x ## do stuff with q, because otherwise you're dereferencing self.x all the damn time
- mathisfun123 7mo ago> it is nuts that in an object method, there is a performance enhancement through caching a member value i don't understand what you think is nuts about this. it's an interpreted language and the word `self` is not special in any way (it's just convention - you can call the first param to a method anything you want). so there's no way for the interpreter/compiler/runtime to know you're accessing a field of the class itself (let alone that that field isn't a computed property or something like that). lots of hottakes that people have (like this one) are rooted in just a fundamental misunderstanding of the language and programming languages in general <shrugs>.
- pjmlp 7mo agoGreat to see this going, Python also deserves a JIT, and given that only few bother with PyPy or GraalPy, shipping into the CPYthon is the only way to have less "rewrite into XYZ". Kudos to those involved into making it happen.
- openclaw01 7mo ago[dead]
- openclaw01 7mo ago[dead]
- qy-mj 7mo ago[flagged]
- seanw444 7mo agoJumping 12 major versions to one that doesn't exist yet must yield quite the performance boost.
- 12_throw_away 7mo agoAfter installing all released versions of python locally, I can confirm that the `python3.15` command is not only very fast, but is guaranteed not to diverge! $ time python3.15 -c "while True: print('hello world')" bash: python3.15: command not found real 0m0.005s user 0m0.000s sys 0m0.002s
- a3w 7mo agoOver 100% speedup sound like "the code compiled before you asked the compiler to start working". `from future import time_travel`
- quietbritishjim 7mo agoIf the speed of a car increases by 100% does that mean that it arrives at its destination before it left? No, it just means it took 50% of the time it would have otherwise. But I do agree that it would be a bit clearer to talk in terms of time taken rather than speedup % i.e. instead of "20% slowdown to over 100% speedup" it's clearer to say "takes between 50% and 125% of the original time". (Especially since people very often say things like "3 times faster", which technically means 4 times as fast, when they should say "3 times as fast"; "takes 1/3 of the time" is unambiguous.)