4 ms·
> [...] pypy, a tracing jit for python2 and python3, is another project which gets compiled to C. That's interesting. How does it work? I think most JIT compil
by frankpf 9y ago
> [...] pypy, a tracing jit for python2 and python3, is another project which gets compiled to C.
That's interesting. How does it work? I think most JIT compilers emit assembly directly and then execute it. Does PyPy generate C code while your program is running?
- sinistersnare 9y agoTo add to the other answer, PyPy is a "Meta-JIT compiler", because it generates JITs, it isnt itself a JIT.
- widdma 9y agoPyPy is written in RPython, which can be compiled to C. PyPy itself emits machine instructions like a regular JIT interpreter.
- ufo 9y agoYou might want to read the PyPy's documentation wiki[1] and the Tracing the Meta-Level paper from the PyPy authors[2]. Pypy consists of a Python interpreter written in the RPython language. RPython is a statically typed language that is easy for an opimizing compilation framework to work with. It also happens to be a valid subset of Python but that isn't actually that important. The RPython translation framework has three ways to run RPython code: * The first way is by using an ahead-of-time compiler to convert the RPython to C and then passing that to a C compiler (gcc or clang). * The second way is by compiling the RPython down to RPython bytecode and running that through an RPython bytecode interpreter. (The RPython bytecode interpreter is written in C) * The third way is the trace compiler. It takes in a linearized trace of the program execution (as observed by the RPython bytecode interpreter) and generates optimized machine code for this trace. So, going back to PyPy as a concrete example, this is what is going on: PyPy is a Python interpreter written in RPython. The pypy executable contains two versions of this python interpreter, both created by the RPython translation framework. The first version is the one where the rpython code for the python interpreter gets compiled to C. The second version is one where the rpython code for the python interpreter is converted to rpython bytecode (which is stored in the data section of the pypy binary) and where this bytecode is in turn executed by the rpython bytecode interpreter. When you run a python program with pypy, it starts being executed by the first interpreter. When this interpreter detects a hot loop in the python program, it transfers control to the RPython bytecode interpreter. This bytecode interpreter executes for one iteration of the loop, and records a trace of what rpython bytecode instructions were executed in this iteration. Then, it uses the jit-compiler to directly generate machine code for this trace, and transfers execution to that. If it goes according to plan, the base C interpreter will have speed comparable to CPython (or slightly slower), the tracing interpreter will be very slow (but for just one iteration) and the machine code for the traces will be very fast. ------- You might be wondering: why bother with so many different steps? 1) Why write a python interpreter in rpython and then run that on an rpython jit compiler instead of just writing a JIT compiler for Python? The PyPy devs had already tried that with the Psyco JIT compiler. 2) Why generate C code for the first interpreter if you already have a JIT compiler that can generate machine code? The JIT compiler can only compile linear traces, and can't compile whole functions. 3) Why not use C as an intermediate language for the JIT compiler as well? While C is a workable intermediate language for method-at-a-time compilation, it is less suited for trace compilers.