3 ms·
> every memory access can cause a synchronous signal (SIGSEGV/SIGBUS) AFAIU you have all the same problems as soon as you translate more than one instruction a
by aleclm 3y ago
> every memory access can cause a synchronous signal (SIGSEGV/SIGBUS)
AFAIU you have all the same problems as soon as you translate more than one instruction at a time and allow merging them, which QEMU does, even if just at basic block level. IIRC post-SIGSEGV state is not 100% correct in QEMU, but I'd need to investigate some more.
We built a dynamic binary translator for a customer dealing with mainframes using ORC JIT where we translate large amounts of code, handle self-modifying code and caching things on disk (similar approach to the one you suggest). Compile time is a problem, but then we the performance results are impressive. Giving visibility over loops to LLVM helps a lot.
But the real reason why nothing like this is upstream is that it's difficult to get things upstream, it requires a lot of effort compared to putting together a PoC. :)
- aengelke 3y ago> allow merging them, which QEMU does, even if just at basic block level. IIRC post-SIGSEGV state is not 100% correct in QEMU A quick and very limited testing shows that all registers and even status flags are correct in the ucontext in the signal handler. Glancing at the source code, TCG optimizations (e.g., liveness analysis) primarily apply to temporaries, but the architectural registers are always updated. That's also why QEMU is so slow (and easy to beat in papers, which very often disregard strict correctness). Function-level lifting to LLVM gives massive performance improvements, but sacrifices correctness w.r.t. signals (synchronous and also asynchronous, unless you add a check for pending signals to every loop, or somehow else recover the state in not-too-distant time).
- aleclm 3y ago> architectural registers are always updated In tiny code, the guest registers (global TCG variables) are stored in the host's registers until you either call an helper which can access the CPU state or you return (`git grep la_global_sync`). This is the reason why QEMU is not so terribly slow. But after a check, this also happens when you access the guest memory address space! https://github.com/qemu/qemu/blob/master/include/tcg/tcg-opc.h#L203C16-L203C19 https://github.com/qemu/qemu/blob/master/include/tcg/tcg-opc... (TCG_OPF_SIDE_EFFECTS is what matters) But still, in the end, it's the same problem. What QEMU does, can be done in LLVM too. You could probably be more efficient in LLVM by using the exception handling mechanism (invoke and friends) to only serialize back to memory when there's an actual exception, at the cost of higher register pressure. More or less what we do here: https://rev.ng/downloads/bar-2019-paper.pdf https://rev.ng/downloads/bar-2019-paper.pdf