11 ms·
Reflections on 2 years of CPython's JIT Compiler
- ggm 1y agoWhat fundamentals would make the jit, this specific jit faster? Because if it's demonstrably slower, it begs the question if it can be faster or is inherently slower than a decent optimisation path through a compiler. At this point it's a great didactic tool and a passion project surely? Or, has advantages in other dimensions like runtime size, debugging, and .pyc coverage, or in thread safe code or ...
- teruakohatu 1y agoThe article points out they have only begun adding optimisers to the jit compiler. Unoptimised jit < optimised interpreter (at least in this instance) They are working on it presumably because they think there will eventually be a speed ups in general or at least for certain popular workloads.
- taeric 1y agoThe article also specifically calls out machine code generation as a separate thing. I confess that somewhat surprises me, as I would expect getting machine code generated would be a main source of speed up for a JIT? That and counter based choices on what optimizations to perform? Still, to directly answer the first question, I would hope even if there wasn't obvious performance improvements immediately, if folks want to work on this, I see no reason not to explore it. If we are lucky, we find improvements we didn't expect.
- adrian17 1y ago> I confess that somewhat surprises me, as I would expect getting machine code generated would be a main source of speed up for a JIT? My understanding is that the basic copy-and-patch approach without any other optimizations doesn’t actually give that much. The difference between an interpreter running opcodes A,B,C and a JIT emitting machine code for opcode sequence A,B,C is very little - the CPU running the code will execute roughly the same instructions for both, the only difference is that the jit avoids doing an op dispatch between each op - but that’s already not that expensive due to jump threading in the interpreter. Meanwhile the JIT adds an extra possible cost of more work if you ever need to jump from JIT back to fallback interpreter. But what the JIT allows is to codegen machine code corresponding to more specialized ops that wouldn’t be that beneficial in the interpreter (as more and smaller ops make it much worse for icaches and branch predictors). For example standard CPython interpreter ops do very frequent refcount updates, while the JIT can relatively easily remove some sequences of refcount increments followed by immediate decrements in the next op. Or maybe I misunderstood the question, then in other words: in principle copy-and-patch’s code generation is quite simple, and the true benefits come from the optimized opcode stream that you feed it that wouldn’t have been as good for the interpreter.
- taeric 1y agoRight, that is basically what I was asking. Essentially, I expected the machine code to be a bit of an unrolling of the interpreter over the opcodes that a piece of code is executing. That my intuition is wrong here doesn't shock me, I should add. It was still a surprise and it will get me to update my idea on what the interpreter is doing.
- moregrist 1y agoA byte code interpreter is, very approximately, a lookup table of byte code instructions that dispatches each instruction to highly optimized assembly. This will almost certainly outperform a straight translation to poorly optimized machine code. Compilers are structured in conceptual (and sometimes distinct) layers. In a classic statically-typed language will only compile-time optimizations, the compiler front-end will parse the language into a abstract syntax tree (AST) via a parse tree or directly, and then convert the AST into the first of what may be several intermediate representations (IRs). This is where a lot of optimization is done. Finally the last IR is lowered to assembly, which includes register allocation and some other (peephole) optimization techniques. This is separate from the IT manipulation so you don’t have to write separate optimizers for different architectures. There are aspects of a tracing JIT compiler that are quite different, but it will still use IR layers to optimize and have architecture-dependent layers for generating machine code.
- taeric 1y agoRight, I guess my main surprise is that the PyPy byte code interpreter is as fast as it is. My understanding is obviously outdated on how it is implemented; but I thought its claim to fame was that it was purely written in python. I'm assuming the subset of python it is implemented in is fairly restricted? That or my understanding was wrong in other ways. :D
- pjmlp 1y agoYes, it is called RPython, https://rpython.readthedocs.io/en/latest/getting-started.html https://rpython.readthedocs.io/en/latest/getting-started.htm...
- MobiusHorizons 1y agoThe way I understand it, the machine code generator emits machine code for some particular piece of bytecode (or whatever the JIT IR is). This is almost like an assembler and probably has templates that it expands. It is important for this machine code to be fast, but it each template is at a pretty low level, and lacks the context for structural optimizations. The optimizer works at a higher level of abstraction, and can make these structural optimizations. You can get very large speed-ups when you can remove code that isn't necessary, or emit equivalent code that has a lower complexity or memory overhead. Typical examples of things optimizers do are * use registers instead of memory for function arguments * constant folding * function inlining * loop unrolling I don't know if that's exactly how it works for this particular effort, but that would be my expectation.
- pizlonator 1y agoIn JavaScript, an unoptimizing JIT (no regalloc, no optimizations that look at patterns of ops, no analysis) is faster than the interpreter because it eliminates opcode dispatch. Adding more optimizations improves things from there. But the point is, a JIT can be a speedup just because it isn’t an interpreter (it doesn’t dynamically dispatch ops).
- deleted 1y ago[deleted]
- eigenspace 1y agoIt turns out that if you have language semantics that make optimizations hard, making a fast optimizing compiler is hard. Who woulda thunk? To be clear, this seems like a cool project and I dont want to be too negative about it, but i just think this was an entirely foreseeable outcome, and the amount of people excited about this JIT project when it was announced shows how poorly a lot of people understand what goes into making a language fast.
- almostgotcaught 1y ago> It turns out that if you have language semantics that make optimizations hard, making a fast optimizing compiler is hard. Who woulda thunk? Is this in the article? I don't see Python's semantics mentioned anywhere as a symptom (but I only skimmed). > shows how poorly a lot of people understand what goes into making a language fast. ...I'm sorry but are you sure you're not one of these people? Some facts: 1. JS is just as dynamic and spaghetti as Python and I hope we're all aware that it has some of the best jits out there; 2. Conversely, C++ has many "optimizing compiler[s]" and they're not all magically great by virtue of compiling a statically typed, rigid language like C++.
- o11c 1y agoJS is absolutely not as dynamic as Python. It supports `const`ness, and uses it by default for classes and functions.
- throwaway032023 1y agoI remember when pypy was only 25x slower than c python.
- bgwalter 1y agoAccording to the promises of the Faster CPython Team, the JIT with a >50% speedup should have happened two years ago. Everyone knows Python is hard to optimize, that's why Mojo also gave up on generality. These claimed 20-30% speedups, apparently made by one of the chief liars who canceled Tim Peters, are not worth it. Please leave Python alone.
- deleted 1y ago[deleted]
- notatallshaw 1y agoTwo years ago was Python 3.11, my real world workloads did see a ~15-20% improvement in performance with that release. I don't remember the Faster CPython Team claiming JIT with a >50% speedup should have happened two years ago, can you provide a source? I do remember Mark Shannon proposed an aggressive timeline for improving performance, but I don't remember him attributing it to a JIT, and also the Faster CPython Team didn't exist when that was proposed. > apparently made by one of the chief liars who canceled Tim Peters Tim Peters still regularly posts on DPO so calling him "cancelled" is a choice: https://discuss.python.org/u/tim.one/activity https://discuss.python.org/u/tim.one/activity. Also, I really can not think who you would be referring to as part of the Faster CPython Team, of which all the former members I am aware of largely stayed out of the discussions on DPO.
- ecshafer 1y agoDoes anyone know why for example the Ruby team is able to create JITs that are performant with comparative ease to Python? They are in many ways similar languages, but Python has 10x the developers at this point.
- cuchoi 1y agoFunding? Seems like the development was funded by Shopify and they got a ~20% performance improvement. https://shopify.engineering/ruby-yjit-is-production-ready https://shopify.engineering/ruby-yjit-is-production-ready A similar experience in the Python community is that Microsoft funded "Faster CPython" and they made Python 20-40% faster.
- ecshafer 1y agoThe funding is one angle, but the Shopify Ruby team isn't that big (<10 people iirc). Python is used extensively at just about every tech company, and Meta, Apple, Microsoft, Alphabet, and Amazon each have at least 10x as many engineers as Shopify. This makes me think that there must be some kind of language/ecosystem reason that makes Python much harder than Ruby to optimize.
- UncleEntity 1y agoProbably the methods they use as well. I may not be completely accurate on this because there's not a whole lot of information on how Python is doing their thing so... The way (I believe) Python is doing it is to take code templates and stitching them together (copy & patch compilation) to create an executable chunk of code. If, for example, one were to take the py-bytecode and just stitch all the code chunks together all you can realistically expect to save is the instruction dispatch operations, which the compiler should make really fast anyway, which leaves you at parity with the interpreter since each code chunk is inherently independent so the compiler can't do its magic on the entire code chunk. Basically this is just inlining the bytecode operations. To make a JIT compiler really excel you'd need to do something like take all the individual operations of each individual opcode and lower that to an IR and then optimize over the entire method using all the bells and whistles of modern compilers. As you can imagine this is a lot more work than 'hacking' the compiler into producing code fragments which can be patched together. Modern compilers are really good at these sorts of things and people have been trying to make the Python interpreter loop as efficient as possible for a long time so there's a big hurdle to overcome here. I've (or more accurately, Claude) has been writing a bytecode VM and the dispatch loop is basically just a pointer dereference and a function call which is about as fast as you can get. Ok, theoretically, this is how it works as there's also a check to make sure the opcode is within range as the compiler part is still being worked on and it's good for debugging but foundationally this is how it works. From what I've gleaned from the literature the real key to making something like copy & patch work is super-instructions. You take common patterns, like MULT+ADD, and mash them together so the C compiler can do its magic. This was maybe mentioned in the copy & patch paper or, perhaps, they only talked about specialization based on types, don't actually remember. So, yeah, if you were just competing against a basic tree-walking interpreter then copy & patch would blow it out of the water but C compilers and the Python interpreter have both had million of people hours put into them so that's really tough competition.
- firesteelrain 1y agoWe have had really good success using Cython which makes many calls into the CPython interpreter and CPython Standard Libraries.
- serjester 1y agoThis article doesn't do the best job explaining the broader picture - stability has been their number one priority up to this point. - Most of the work has just been plumbing. Int/float unboxing, smarter register allocation, free-threaded safety land in 3.15+. - Most JIT optimizations are currently off by default or only triggers after a few thousand hits, and skips any byte-codes that look risky (profiling hooks, rare ops, etc.). I really recommend this talk with one of the Microsoft faster Cpython developers for more details, https://www.youtube.com/watch?v=abNY_RcO-BU https://www.youtube.com/watch?v=abNY_RcO-BU
- kenjin4096 1y agoHi, author of the post here, stability indeed has been a priority. There are some points which are not exactly the case though: > - Most of the work has just been plumbing. Int/float unboxing, smarter register allocation, free-threaded safety land in 3.15+. The first part is true, but for the second sentence: none of that is guaranteed to land in 3.15+. We proposed to land them, that doesn't mean they will. Landing a PR in CPython is subject to maintainer time and reviewer approval, which doesn't always happen. I proposed a few optimizations for 3.14 that never landed. > Most JIT optimizations are currently off by default or only triggers after a few thousand hits It is indeed true we only trigger after a few thousand hits, but all optimizations that we currently have are always enabled. We don't sandbag the JIT on purpose.
- gjvc 1y agonot so long ago some people were saying that pypy should be the de-facto reference implementation because of its speed
- pizlonator 1y agoJIT and VM writer here. I’m also pretty clued in on how CPython works because I ported it to Fil-C. I think if I was being paid to make CPython faster I’d spend at least a year changing how objects work internally. The object model innards are simply too heavy as it stands. Therefore, eliminating the kinds of overheads that JITs eliminate (the opcode dispatch, mainly) won’t help since that isn’t the thing the CPU spends much time on when running CPython (or so I would bet).
- cs_throwaway 1y agoDo you think it may be feasible to do this and maintain the FFI?
- pizlonator 1y agoThat's the hard part! I think that the FFI makes it super hard to do most of the optimizations I'd want to do. Maybe it makes them impossible even. The game is to find any chance for size reduction and fast path simplification that doesn't upset FFI
- kzrdude 1y agoMany changes of that kind have been made by the faster-cpython team I believe, Mark Shannon was rather focused on it (and had a decade of experience of that kind of tweaks to python). But I'm trying to find/recall a blog post that detailed the different steps in shrinking the CPython object struct... If you say that's not enough, more radical changes needed, I would understand.
- pizlonator 1y agoGotta be careful about the tone and mindset here. Given any arbitrarily optimized thing, it is always possible to optimize it more. And the fact that it's possible to optimize it more is not meant as a criticism of folks who did the previous optimizations. So, I have no doubt that Mark and others have worked on exactly the thing I'm talking about and that they've gotten wins. And I have no doubt that more can be done. Also, not saying I would do a better job at it than Mark or anyone else