6 ms·
Show HN: I wrote a WebAssembly Interpreter and Toolkit in C
- 4984 4y agoI made Web49 because there are not many good tools for WebAssembly out there. WABT is close, but the interpreter is too slow and the tools megabytes in size each. Wasm3 is a bit faster but only contains an interpreter, nothing else. Tooling for WebAssembly is held mostly by the browser vendors. It is such a nice format to work with when one removes all the fluff. WebAssembly tooling should not take seconds to do what should take milliseconds, and it should be able to be used as a library, not just a command line program. I developed a unique way to write interpreters based on threaded code jumps and basic block versioning when I made MiniVM (https://github.com/FastVM/minivm https://github.com/FastVM/minivm). It was both larger and more dynamic than WebAssembly. Web49 started as a way to compile WebAssembly to MiniVM, but soon pivoted into its own Interpreter and tooling. I could not be happier with it in its current form and am excited to see what else It can do, with more work.
- haberman 4y ago> I developed a unique way to write interpreters based on threaded code jumps and basic block versioning when I made MiniVM (https://github.com/FastVM/minivm https://github.com/FastVM/minivm). It was both larger and more dynamic than WebAssembly. I'd be very interested to read more about this. It looks like you are using "one big function" with computed goto (https://github.com/FastVM/Web49/blob/main/src/interp/interp.c#L573-L586 https://github.com/FastVM/Web49/blob/main/src/interp/interp....). My experience working on this problem led me to the same conclusion as Mike Pall, which is that compilers do not do well with this pattern (particularly when it comes to register allocation): http://lua-users.org/lists/lua-l/2011-02/msg00742.html http://lua-users.org/lists/lua-l/2011-02/msg00742.html I'm curious how you worked around the problem of poor register allocation in the compiler. I've come to the conclusion that tail calls are the best solution to this problem: https://blog.reverberate.org/2021/04/21/musttail-efficient-interpreters.html https://blog.reverberate.org/2021/04/21/musttail-efficient-i...
- naasking 4y ago> My experience working on this problem led me to the same conclusion as Mike Pall, which is that compilers do not do well with this pattern Note that that message is from twelve years ago. A lot's changed since then, not just in compilers but in CPUs. Branch prediction is a lot better now.
- haberman 4y agoMike's primary complaint is bad register allocation. It is very important to keep the most important state consistently in registers. In my experience, compilers still struggle to do good register allocation in big and branchy functions. Even perfect branch prediction cannot solve the problem of unnecessary spills.
- 10000truths 4y agoDoes providing a hint to the compiler using the register keyword address the issue sufficiently?
- haberman 4y agoNo, most compilers ignore the register keyword, see: https://stackoverflow.com/a/10675111 https://stackoverflow.com/a/10675111
- JonChesterfield 4y agoNearly. You need register and to also pass them into (potentially no-op) inline asm. `register int v("eax")` iirc, but it's been years since I did this. The 'register' is indeed largely ignored, but it has the additional somewhat documented meaning of 'when this variable goes into inline asm, it needs to be in that register'. In between asm blocks it can be elsewhere - stack or whatever - but it still gives the regalloc a really clear guide to work from.
- lifthrasiir 4y ago
- deleted 4y ago[deleted]
- Octokiddie 4y agoAny ideas on why miniwasm performs better on all the benchmarks except "trap," on which it performs decidedly worse?
- 4984 4y agoThe benchmarks were run on MacOS, and actually execute an interrupt for debugging, MacOS then checks if the process is being debugged. Wasm3 just exit(1) and prints a message. And as to why the rest are faster, I spent much time optimizing the interpreter and learning what the best way to write interpreters is. Its mostly jump threading and Mixed Data.
- titzer 4y agoI found that most Wasm interpreters are not particularly good at calls. Wizard is not as fast as wasm3 or wamr in raw speed, but is much faster on calls, particularly because it does not copy arguments (value stacks can be overlapped). But Wizard's primary motivation is to be memory efficient, so it interprets in-place. It also supports instrumentation. Nice work!
- fwsgonzo 4y agoDon't take this as anything other than speculation: I wonder if wasm3 is using musttail with opaque function calls in the instruction handlers. It will demolish performance, which is why I am only using computed gotos in mine (when available). Even switch-case is faster than musttail when you have to leave the tco-jumps. Which is (as an example) why one should not measure performance by fibonacci number generation. :)
- haberman 4y ago> I wonder if wasm3 is using musttail with opaque function calls in the instruction handlers. It will demolish performance, which is why I am only using computed gotos in mine (when available). Even switch-case is faster than musttail when you have to leave the tco-jumps. This doesn't match with my experience. After working on this problem a lot, I came to the conclusion that musttail with opaque function calls is one of the best ways of getting good code out of the compiler: https://blog.reverberate.org/2021/04/21/musttail-efficient-interpreters.html https://blog.reverberate.org/2021/04/21/musttail-efficient-i...
- duped 4y agoHow does it compare to Wasmtime? Do you have a list of supported wasm features?
- habibur 4y agoI am afraid it doesn't work on linux.
- heleninboodler 4y agoWorked flawlessly for me on linux. The readme should mention 'emcc' and 'wasm3' as prerequisites for following the bench.py instructions, though.
- thechao 4y agoI'm on macOS, and if I do this: > git clone ... > make -j > ./bin/wasm2wat ./test/core/address.wast I get: >>> ./bin/wat2wasm test/core/address.wast unexpected word: `` byte=256
- muricula 4y agoHas this been fuzz tested? Fuzz testing is one of the best ways of discovering security vulnerabilities in C code, and tools like libfuzzer and asan or AFL make it easier than ever.
- vmafficianado 4y agoImpressive! Will work stop on MiniVM as result of Web49?
- mouse_ 4y agoI'd love to see a benchmark of this vs. libwasm https://github.com/SerenityOS/serenity/tree/master/Userland/Libraries/LibWasm https://github.com/SerenityOS/serenity/tree/master/Userland/...
- cornstalks 4y agoCan this share memory with the host? wasm3 doesn't allow this[1] and requires you to allocate VM memory and pass it to the host, but that has several downsides (some buffers come from external sources so this requires a memcpy; the VM memory location isn't stable so you can't store a pointer to it on the host; etc.). I'm really interested in a fast interpreter-only Wasm VM that can allow the host to share some of its memory with the VM. [1]: https://github.com/wasm3/wasm3/issues/114 https://github.com/wasm3/wasm3/issues/114