9 ms·
Wasm3 – A high performance WebAssembly interpreter in C
- setheron 7y agoThe neater article seems to be about M3 interpreter https://github.com/soundandform/m3#m3-massey-meta-machine https://github.com/soundandform/m3#m3-massey-meta-machine Tbh, I couldn't get the eureka moment though. Might try to read in the AM ;)
- thermals 7y agoYeah, this is a good way to design a fast interpreter! It's traditionally called a "threaded interpreter", or (somewhat confusingly) "threaded code": https://en.wikipedia.org/wiki/Threaded_code https://en.wikipedia.org/wiki/Threaded_code http://www.complang.tuwien.ac.at/forth/threaded-code.html http://www.complang.tuwien.ac.at/forth/threaded-code.html You can see an example of this particular implementation style (where each operation is a tail call to a C function, passing the registers as arguments) at the second link above, under "continuation-passing style". One of the big advantages of a threaded interpreter is relatively good branch prediction. A simple switch-based dispatch loop has a single indirect jump at its core, which is almost entirely unpredictable -- whereas threaded dispatch puts a copy of that indirect jump at the end of each opcode's implementation, giving the branch predictor way more data to work with. Effectively, you're letting it use the current opcode to help predict the next opcode!
- vshymanskyy 7y agoYeah, but... It's not only the "threaded code" approach, that makes Wasm3 so fast. In fact, Intel's WAMR also utilizes this method, yet is 30 times slower..
- sound_and_form 7y agoThanks for the links! I've long searched google trying to find a similar "tail-cail" interpreter. No wonder I couldn't hit anything -- it was so poorly named! :)
- 29athrowaway 7y agoThe motivation for WebAssembly rather than plain ARM or x86 assembly = portability, security. It would be interesting to see how this is designed for security in mind.
- kick 7y agoThis is designed for microcontrollers. If you're running untrusted code on microcontrollers, you've got bigger problems.
- saagarjha 7y agoI wouldn't go as far as to say it's designed for microcontrollers; as others have mentioned, there's a number of potential applications for this in other contexts as well.
- DarthGhandi 7y agoPardon my wasm illiteracy here, what exactly makes it more secure? Struggling to see it.
- earenndil 7y agoSandboxing is trivial, because you can cover all paths to the outside world.
- DarthGhandi 7y agoI edited my earlier comment to remove the sandboxing part thinking this may be embedded specific, so not sure you saw that, so ask again how is this special? would you rather inspect potentially malicious wasm or js?
- 29athrowaway 7y agoAll web APIs are designed to be secure in the context of the web. Needs to be memory safe otherwise a wasm program can execute arbitrary code, access memory that it should not, etc.
- CharlesW 7y agoWhy is an interpreter desirable when JIT compilers create significantly faster code? Is this primarily about embedded use?
- Koshkin 7y agoOn the embedded side, I'd rather see a hardware interpreter.
- leetrout 7y agoI'm hardware ignorant... would that be a system on a chip?
- newnewpdro 7y agoI suspect they're just saying they'd rather see a CPU implementing the wasm isa, in a facetious way.
- sebcat 7y agoAVR32 with JEM does this to some extent for Java. http://ww1.microchip.com/downloads/en/devicedoc/doc32000.pdf http://ww1.microchip.com/downloads/en/devicedoc/doc32000.pdf 3. Java Extension Module The AVR32 architecture can optionally support execution of Java bytecodes by including a JavaExtension Module (JEM). This support is included with minimal hardware overhead. EDIT: Not by implementing the ISA verbatim, but still neat though
- saagarjha 7y agoARM tried to do this with Jazelle, but the effort floundered: https://en.wikipedia.org/wiki/Jazelle https://en.wikipedia.org/wiki/Jazelle
- MaxBarraclough 7y agoIt doesn't make much sense to build hardware to run an intermediate language. Java bytecode isn't optimised; almost all optimisation is meant to happen in the JIT. If you build a processor that runs bytecode directly, when does the optimisation occur? Presumably never. Low-horsepower platforms are probably the best place to give it a go, as they may struggle to run a respectable JIT compiler, but as you say, Jazelle didn't catch on.
- ridiculous_fish 7y agoThis is pretty exciting if real: > Bytecode/opcodes are translated into more efficient "operations" during a compilation pass, generating pages of meta-machine code WASM compiled to a novel bytecode format aimed at efficient interpretation. > Commonly occurring sequences of operations can can also be optimized into a "fused" operation. Peephole optimizations producing fused opcodes, makes sense. > In M3/Wasm, the stack machine model is translated into a more direct and efficient "register file" approach WASM translated to register-based bytecode. That's awesome! > Since operations all have a standardized signature and arguments are tail-call passed through to the next, the M3 "virtual" machine registers end up mapping directly to real CPU registers. This is some black magic, if it works!
- kevingadd 7y agoIR getting converted into an interpreter-oriented bytecode is pretty common. Mono does it for its interpreter and IIRC, Spidermonkey has historically done that as well. I'm not sure if V8 has ever interpreted from an IR (maybe now?) but you could view their original 'baseline JIT' model as converting into an interpreter-focused IR, where the IR just happened to be extremely unoptimized x86 assembly. Translating the stack machine into registers was always a core part of the model but it's interesting to me that even interpreters are doing it. The necessity of doing coloring to assign registers efficiently is kind of unfortunate, I feel like the WASM compiler would have been the right place to do this offline.
- jashmatthews 7y ago> The necessity of doing coloring to assign registers efficiently is kind of unfortunate Register based VMs like Lua don't do this. The register allocation is incredibly simple https://github.com/LuaJIT/LuaJIT/blob/v2.1/src/lj_parse.c#L370 https://github.com/LuaJIT/LuaJIT/blob/v2.1/src/lj_parse.c#L3...
- tom_mellior 7y agoBut that's an allocator to virtual registers that don't try to correspond to (a valid number of) physical CPU registers. Sure it's easy to allocate to a large number of registers. It's harder to do it to a small number, like the project discussed here seems to claim to do.
- haberman 7y agoThese are impressive performance numbers. > Because operations end with a call to the next function, the C compiler will tail-call optimize most operations. It appears that this relies on tail-call optimization to avoid overflowing the stack. Unfortunately this means you probably can't run it in debug mode.
- vshymanskyy 7y agoIt's not that bad even in debug mode (or without TCO). Just not optimal. Also, there is a way to rework this part, so it does not rely on compiler TCO.
- haberman 7y agoIf the jump to the next opcode is a tail call, wouldn't an arbitrarily long sequence of instructions take arbitrarily much stack space?
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- jononor 7y agoImpressive list of constrained targets for embedded. The AtMega1284 microcontroller for example has only 16 KB of RAM. Which is a lot for an 8-bit micro, but pretty standard for a modern application processors.
- vshymanskyy 7y agoYup. TinyBLE is nRF51 SoC with 16Kb SRAM as well.
- MuffinFlavored 7y ago> Node v13.0.1 (interpreter) 28 59.5x https://github.com/wasm3/wasm3/blob/master/test/benchmark/coremark/README.md https://github.com/wasm3/wasm3/blob/master/test/benchmark/co... 59.5x faster than node.js at what? Executing WebAssembly?
- vshymanskyy 7y agoV8 has a built-in (pure, no JIT etc.) interpreter of WASM. Which is quite slow according to this test.