3 ms·
> I'd be very eager to see the CPython benchmarks! In the talk on youtube, the author mentions that it’s not faster than mainline CPython yet (it is slightly
by adrian17 3y ago
> I'd be very eager to see the CPython benchmarks!
In the talk on youtube, the author mentions that it’s not faster than mainline CPython yet (it is slightly faster than experimental off-by-default microoperation support it’s built on top of, but it was already slower than mainline, so it cancels out at best). I think the idea is for it to be merged, but only enabled by default once it becomes worth it; and that’s why the perf numbers aren’t advertised yet.
Still, I wonder what the expected peak improvement is. Looking at the current generated assembly, there’s definitely room to improve, but there’s only so much one can do without touching the data model.
- lifthrasiir 3y agoThe goal is to enable JIT codegen without sacrificing too much performance and adding too much maintenance burden, and a functional JIT implementation needs a few more components other than that---most notably a facility to monitor and trace function calls for the eventual JIT compilation. Consider the OP to be one of intermediate goals, not the eventual goal.
- adrian17 3y agoI don't think we disagree that the long-term goal is to _eventually_ make it faster :) I rather meant to temper the enthusiasm that some could have upon seeing "JIT" and immediately trying to compare with, say, PyPy. > enable JIT codegen without sacrificing too much performance This is the part I don't buy. The main point of a JIT is performance, so by definition I don't see it being enabled unless it improves performance across the board. What I wonder is if the current approach, stated as "copy-and-patch auto-generated code for each opcode", can ever reach that point without being replaced by a completely different design along the way. AFAIK, as is, the main difference between running the interpreter loop composed of normally compiled opcodes and JIT copy-and-patching these opcodes is lack of the opcode dispatch logic running between each op - which is good, but also countered by slightly worse quality of the copied code.
- lifthrasiir 3y ago> What I wonder is if the current approach, stated as "copy-and-patch auto-generated code for each opcode", can ever reach that point without being replaced by a completely different design along the way. Of course this approach produces a worse code than a full compiler by definition---stencils would be too rigid to be further optimized. A stencil conceptually maps to a single opcode, so the only way to break out of this restriction is to add more opcodes. And there are only so many opcodes and stencils you can prefare. But I think you are thinking too much about a possibility to make Python as fast as, say, C for at least some cases. I believe that it won't happen at all, and the current approach clearly points why. Let's consider a simple CPython opcode named `BINARY_ADD` which has a stack effect of `(a b -- sum)`. Ideally it should eventually be compiled down to a fully specialized machine code something like `add rax, r12`, plus some guards. But the actual implementation (`PyNumber_Add` [1]) is far more complex: it may call at most 3 "slot" calls that may add or concatenate arguments, some of them may call back to a Python code. So let's assume that we have done type specialization and arguments are known to be integers. That will result in a single slot call to `PyLong_Add` [2], which again is still complex because CPython has two integer representations. Even when they are both "compact", i.e. at most 31/63 bits long, it may still have to switch to another representation when the resulting sum is no longer compact. So a fully specialized machine code would be only possible when both arguments are known to be integers and compact and have one more spare bit to prevent an overflow. That sounds way more restrictive. [1] https://github.com/python/cpython/blob/36adc79041f4d2764e1daf7db5bb478923e89a1f/Objects/abstract.c#L904-L971 https://github.com/python/cpython/blob/36adc79041f4d2764e1da... [2] https://github.com/python/cpython/blob/36adc79041f4d2764e1daf7db5bb478923e89a1f/Objects/longobject.c#L3440-L3470 https://github.com/python/cpython/blob/36adc79041f4d2764e1da... An uncomfortable truth is that all these explanations also almost perfectly apply to JavaScript---the slot resolution would be the `[[ToNumber]]` internal function and multiple representations will be something like V8's Smi. Modern JS engines do exploit most of them, but at the expense of extremely large codebase with tons of potential attack surfaces. It is really expensive to maintain, and people don't really realize that no performant JS engine was ever developed by a small group of developers. You have to cut some corners. In comparison, CPython's approach is essentially inside out. Any JIT implementation will require you to split all those subtasks into small bits that can be either optimized out or baked into a generated machine code. So what if we start with subtasks without thinking about JIT in the first place? This is what a specializing adaptive interpreter [3] did. The current CPython already has two tiers of interpreters, and micro-opcodes can only appear in the second tier. With them we can split larger opcodes into smaller ones, possibly with optimizations, but its performance is limited by the dispatch logic. The copy-and-patch JIT is not as powerful, but it does eliminate the dispatch logic without large design changes and it's a good choice for this purpose. In the best scenario, it will eventually hit the limit of what's possible with copy-and-patch and a full compiler will be required at that point. But until that point (which may never come as well), this approach allows for a long time of incremental improvements without disruption. [3] https://peps.python.org/pep-0659/ https://peps.python.org/pep-0659/
- resoluteteeth 3y ago> The goal is to enable JIT codegen without sacrificing too much performance and adding too much maintenance burden, and a functional JIT implementation needs a few more components other than that---most notably a facility to monitor and trace function calls for the eventual JIT compilation. Consider the OP to be one of intermediate goals, not the eventual goal. It seems like the copy and patch approach is sort of somewhere inbetween an interpreter and a traditional JIT, and the authors of the original copy and patch paper seem to be trying to use it to replace things like the baseline compiler in the two-tier baseline/optimizing compiler strategy used for things like webassembly. Because of this, is it really necessary to add tracing and try use a two tier interpeter/copy-and-patch JIT approach for this python JIT? Wouldn't it make more sense to try to get it to be fast enough that the JIT can be used alone?
- lifthrasiir 3y agoSee my other comment for details, but in short, this strategy uses a single code base for both interpreter and JIT. So any further performance improvement will benefit both without any additional work. The traditional JIT-only approach is costly to maintain in comparison.