3 ms·
Thank you for having a look. The reason MiniVM is so benchmark-oriented in its presentation is that MiniVM is ver benchmark-oriented in my workflow. I have tr
by 4984 5y ago
Thank you for having a look.
The reason MiniVM is so benchmark-oriented in its presentation is that MiniVM is ver benchmark-oriented in my workflow.
I have tried to find a faster interpreter without JIT. It is hard to find benchmarks that don't just say that MiniVM is 2-10x faster than the interpreted languages.
- iainmerrick 5y agoYeah, it’s a bit disappointing how slow most non-JIT interpreters are -- I think JIT understandably took the wind out of their sails. The state of the art as I understand it (but maybe you’ve figured out something faster?) is threaded interpreters. I suggest looking at Forth implementations, which is where I think most interpreter innovation has happened.
- isaacimagine 5y agoSo I was talking about threaded code with Shaw the other day, and he gave me a long explanation about the various tradeoffs being made and the resulting techniques used in MiniVM. As I understand it, MiniVM is basically threaded code, with cached opcodes. Essentially, it uses precomputed gotos: // classical computed goto goto *x[y[z]] // MiniVM precomputed goto goto *a[z] This optimization makes MiniVM a tad faster than raw computed gotos, because it's essentially threaded code with some optimizations. He can probably expand on this; if you want to hear the full really interesting explanation yourself, he's pretty active in the Paka Discord[0]. [0]: https://discord.gg/UyvxuC5W5q https://discord.gg/UyvxuC5W5q
- aardvark179 5y agoThat feels like the sort of optimisation that wins in small cases and loses in larger ones. The jump array is going to be 4 or 8 times the size of the opcode array so will put a lot more pressure on the caches as your programs get larger.
- sitkack 5y agoSounds like a great test case for using a tensorflow-lite model to switch between both techniques.
- shuffel 5y agoYears ago I wrote a small toy interpreter based on predereferenced computed gotos with surprisingly good speed. For very small programs without dynamic data allocation churn it could run as fast as 3 times slower than compiled code instead of the expected 10X slower. Good things happen speed-wise when all the byte code and data fits into the L1 cache. Also, branch prediction is near perfect in tight loops.
- MaxBarraclough 5y ago> I suggest looking at Forth implementations, which is where I think most interpreter innovation has happened. Fortunately the Forth community write about their techniques, so you don't have to just eyeball gforth's source-code. A handful of links for those interested: https://www.complang.tuwien.ac.at/papers/ https://www.complang.tuwien.ac.at/papers/ https://github.com/ForthPapersMirror/Papers https://github.com/ForthPapersMirror/Papers https://www.complang.tuwien.ac.at/projects/forth.html https://www.complang.tuwien.ac.at/projects/forth.html https://www.complang.tuwien.ac.at/anton/euroforth/ https://www.complang.tuwien.ac.at/anton/euroforth/ http://soton.mpeforth.com/flag/jfar/vol7.html http://soton.mpeforth.com/flag/jfar/vol7.html http://soton.mpeforth.com/flag/fd/ http://soton.mpeforth.com/flag/fd/ http://www.forth.org/literature.html http://www.forth.org/literature.html http://www.forth.org/fd/contents.html http://www.forth.org/fd/contents.html