4 ms·
I am currently exploring the similar approach. The idea is to compile to llvm IR and then augment it. There's a musttail marker in the language to enforce tail
by rapidlua 6y ago
I am currently exploring the similar approach. The idea is to compile to llvm IR and then augment it. There's a musttail marker in the language to enforce tail calls no matter what the optimisation level is.
Passing VM state via arguments is suboptimal since argument registers are often clobbered and one would normally use callee-saved registers. I have developed a patch for llvm allowing to specify registers used for passing arguments and making more registers available for allocation without spilling.
I intend to reimplement LuaJIT VM in C with IR augmentation to verify the performance.
- fsfod 6y agoAre you gonna use the LuaJIT C VM RaptorJIT is working on https://github.com/raptorjit/raptorjit/pull/254 https://github.com/raptorjit/raptorjit/pull/254 as a base to start from. I'm also big fan of the LuaJIT compiler explorer you made https://luajit.me https://luajit.me
- tomp 6y agoDid you check out the "GHC calling convention"? They implemented it specifically for this purpose. https://llvm.org/docs/LangRef.html#calling-conventions https://llvm.org/docs/LangRef.html#calling-conventions My understanding is, for "fast path" instructions (e.g. integer addition) there's no register spilling necessary, so passing VM state in registers makes sense. For more complex instructions, you rely on LLVM to optimize registers as it would normally, with callee-saved registers. Of course, one must still pass as little state as possible - e.g. just sp, cp, maybe top of stack / heap, and then a pointer to *vm_state for all the rest.
- rapidlua 6y ago> Did you check out the "GHC calling convention"? GHC allocates registers for arguments from the fixed list where most but not all registers are callee-save in other common CC-s. This is a big improvement; still there are valid reasons to want it customisable. E.g. in LuaJIT next instruction is decoded as a part of an instruction handler before dispatching to the next handler. Instruction arguments are passed to a handler; they shouldn't use callee-saved registers. Another complication is the stack frame. There's a requirement to keep the stack pointer aligned (16 byte on x86_64). Call instruction pushes the return address and the pointer becomes misaligned. If the called function wishes to call further functions it has to adjust the stack pointer to make it aligned, even if it doesn't use the stack. The stack pointer is adjusted in a function prologue, hence there's a slowdown even if a nested function call is on a cold path. Whether these micro-optimisations are worthwhile remains to be seen.
- tomp 6y agoI'm not sure I understand your complaint about callee-saved registers. The point is that arguments are passed in registers, as opposed to on the stack, and most calls are tail calls anyways (so `jmp`, not `call`/`ret`). It doesn't matter if the registers are caller-saved or callee-saved, as the context is destroyed anyways, the function never returns. In case of GHC, the functions could return, but the fast-path is that most functions access VM state, and many functions have very little need for registers; therefore, forcing registers be caller-saved would just cause most functions to push and pop the stack without any reason. If non-standard calling convention (e.g. C calling convention) functions are called only rarely, then using caller- or callee-saved registers doesn't really matter.
- rapidlua 6y agoWe are on the same page irt tail calls. I assume that instruction handlers (non-standard calling convention, tail-calls next handler) need to invoke helper functions quite often. These helpers adhere to standard C calling convention. VM state is in registers. Unless they are callee-saved, registers have to be saved to stack and restored after each helper call. (LuaJIT VM uses helper functions heavily.)