4 ms·
You can already do that with QEMU: https://wiki.debian.org/QemuUserEmulation https://wiki.debian.org/QemuUserEmulation However, the code we emit is faster than
by aleclm 3y ago
You can already do that with QEMU: https://wiki.debian.org/QemuUserEmulation https://wiki.debian.org/QemuUserEmulation
However, the code we emit is faster than the one emitted by QEMU.
We have a secret plan to use LLVM in QEMU as a tier 2 JIT engine to optimize hot code paths more aggressively.
This has already been done, but never upstreamed.
Adding an on-disk cache would also be quite useful.
- ris 3y ago> This has already been done, but never upstreamed Assume you're referring to hqemu http://csl.iis.sinica.edu.tw/hqemu/ http://csl.iis.sinica.edu.tw/hqemu/ > Adding an on-disk cache would also be quite useful. +100
- aengelke 3y ago> [LLVM in QEMU] has already been done, but never upstreamed. Yeah, that's a long history, and it never worked great, the architectures are too different. LLVM is great for compiling functions while QEMU compiles basic blocks, which is somewhat required for correctness. LLVM for binary translation requires longer superblocks (HQEMU) or functions (Instrew). I don't see _correct_ binary translation with LLVM as viable anytime soon, though: every memory access can cause a synchronous signal (SIGSEGV/SIGBUS), and I'd expect recovering the full register state at the prior instruction boundary from the LLVM-compiled code to come at severe performance penalties. > Adding an on-disk cache would also be quite useful. I implemented a persistent cache in Instrew to work around LLVM's high compile times. Cache access is done by decoding the code-to-be-translated (very fast) and using a hash of that code for lookup. So shared libraries and even the same code in different binaries will hit the cache, regardless of their address or path name. If translation is fast (QEMU/TCG), caching is probably not very useful, though.
- aleclm 3y ago> every memory access can cause a synchronous signal (SIGSEGV/SIGBUS) AFAIU you have all the same problems as soon as you translate more than one instruction at a time and allow merging them, which QEMU does, even if just at basic block level. IIRC post-SIGSEGV state is not 100% correct in QEMU, but I'd need to investigate some more. We built a dynamic binary translator for a customer dealing with mainframes using ORC JIT where we translate large amounts of code, handle self-modifying code and caching things on disk (similar approach to the one you suggest). Compile time is a problem, but then we the performance results are impressive. Giving visibility over loops to LLVM helps a lot. But the real reason why nothing like this is upstream is that it's difficult to get things upstream, it requires a lot of effort compared to putting together a PoC. :)
- aengelke 3y ago> allow merging them, which QEMU does, even if just at basic block level. IIRC post-SIGSEGV state is not 100% correct in QEMU A quick and very limited testing shows that all registers and even status flags are correct in the ucontext in the signal handler. Glancing at the source code, TCG optimizations (e.g., liveness analysis) primarily apply to temporaries, but the architectural registers are always updated. That's also why QEMU is so slow (and easy to beat in papers, which very often disregard strict correctness). Function-level lifting to LLVM gives massive performance improvements, but sacrifices correctness w.r.t. signals (synchronous and also asynchronous, unless you add a check for pending signals to every loop, or somehow else recover the state in not-too-distant time).
- aleclm 3y ago> architectural registers are always updated In tiny code, the guest registers (global TCG variables) are stored in the host's registers until you either call an helper which can access the CPU state or you return (`git grep la_global_sync`). This is the reason why QEMU is not so terribly slow. But after a check, this also happens when you access the guest memory address space! https://github.com/qemu/qemu/blob/master/include/tcg/tcg-opc.h#L203C16-L203C19 https://github.com/qemu/qemu/blob/master/include/tcg/tcg-opc... (TCG_OPF_SIDE_EFFECTS is what matters) But still, in the end, it's the same problem. What QEMU does, can be done in LLVM too. You could probably be more efficient in LLVM by using the exception handling mechanism (invoke and friends) to only serialize back to memory when there's an actual exception, at the cost of higher register pressure. More or less what we do here: https://rev.ng/downloads/bar-2019-paper.pdf https://rev.ng/downloads/bar-2019-paper.pdf