8 ms·
A First Look at the JIT
- kzrdude 6y agoWith this approach, it shouldn't be too far off to experiment with ahead of time compilation?
- whizzter 6y agoI think there has existed Erlang to C compilers in the past and there are new one popping up every now and then. In general though AOT is in many ways considered a dead end for many unless you have a certain application for it (even if that was my uni thesis subject). You can look back all the way to a 1995 paper (1) (with people that then worked on early Java JIT's) that did a comparison of JIT vs AOT compilation for Self, while good AOT is kind of feasible in many ways there is always situations in languages not initially designed for AOT where a JIT won't be forced to go with conservative estimates that hampers performance. (1) https://www.cs.ucsb.edu/research/tech-reports/1995-04 https://www.cs.ucsb.edu/research/tech-reports/1995-04
- jecel 6y agoThe 1995 paper compares Self 1 and 2 which was a JIT which was based on type inference with Self 3 and 4 which was an adaptive compiler (where Java's Hotspot technology came from) which was able to use type feedback thanks to the second compiler having access to the results of running code from the first compiler. So AOT compilation was not tested in that paper.
- whizzter 6y agoWas a while ago since i read the paper but at the time Agesen's CPA algorithm was among the most powerful for type inference (and while we've had a bunch of incremental improvements in the field there hasn't really been any order of magnitudes better algorithm that can handle edge cases that many dynamic languages produce), so regardless of if the Self runtime was JIT or AOT the algorithm choice showed what was feasible/usable with AOT.
- bcardarella 6y agoThe Lumen project (https://github.com/lumen/lumen https://github.com/lumen/lumen) is building a BEAM bytecode compatible AOT compiler. Its primary compilation target is WASM/WASI. Still under very heavy development. (and some things waiting on updates to WASM spec to land)
- whizzter 6y agoMy take on it is that I like most of the design choices they've made (Been tinkering on a similar dynamic runtime before). With a dynamic language like Erlang (as with JS,Lua,etc) the main benefits won't come from better register allocation,DCE,etc that is the primary reason one would pick LLVM but rather just getting rid of the interpreter overhead (and later the oppurtunity to do some higher level type dependent optimizations that will have a higher impact than register allocation and instruction selection). Type dependant stuff why LLVM is ill suited IMHO and why you only see LLVM being mentioned as the highest optimizataion level by for example JavaScriptCore(Safari) when the runtime hasn't deoptimized code and is pretty certain of the types that will be used. Also they mentioned the complexity of maintaining interpreter/JITed jumps and I'm not surprised since i remember reading some paper about one of their earlier attempts and they were maintaining dual stacks with quite complex cross-boundary jumps. Some might mention that V8 moved away from the always-JIT paradigm to save on memory and startup time but since Erlang is mostly used in servers i think they might see this as a good compromise.
- chrisseaton 6y ago> that will have a higher impact than register allocation and instruction selection The most powerful JIT compiler that I have worked with - Graal - generates code that often on the face of it seems puzzlingly inefficient in terms of registers allocated and instructions selected. Turns out maybe it's another thing that's not as important as all the text books say? The important bit is removing object allocating, removing boxing, deep inlining, and specialisation... when you've done all that work the exact registers and instructions don't seem to really make a huge difference.
- Rochus 6y agoAre there any papers or articles about the mentioned findings?
- chrisseaton 6y agoAbout what findings, sorry? Linear Scan is an example I reach for when I talk about what parts of the compiler are really important, if that's what you mean. http://web.cs.ucla.edu/~palsberg/course/cs132/linearscan.pdf http://web.cs.ucla.edu/~palsberg/course/cs132/linearscan.pdf > The linear scan algorithm is considerably faster than algorithms based on graph coloring, is simple to implement, and results in code that is almost as efficient as that obtained using more complex and time-consuming register allocators based on graph coloring. Also things like escape analysis and inlining are often called 'the mothers of optimisation' because they fundamentally enable so many other optimisations. Not sure there's really a citation for it but I doubt anyone would dispute. I wrote about the impact on optimising Ruby. https://chrisseaton.com/truffleruby/pushing-pixels/ https://chrisseaton.com/truffleruby/pushing-pixels/
- Rochus 6y agoInteresting. Here is also a podcast about the topic: https://thinkingelixir.com/podcast-episodes/017-jit-compiler-for-beam-with-lukas-larsson-and-john-hogberg/ https://thinkingelixir.com/podcast-episodes/017-jit-compiler... Had a brief look at asmjit; as it seems it only supports x86 and x86_64 and is not really an abstraction (i.e. a platform independent IR). I will try to find out why they didn't use e.g. LLVM or sljit. EDIT: according to this article (https://www.erlang-solutions.com/blog/performance-testing-the-jit-compiler-for-the-beam-vm.html https://www.erlang-solutions.com/blog/performance-testing-th...) the speed-up factor caused by the JIT is about 1.3 to 2.3 (as a comparison the speed-up between the PUC Lua 5.1 interpreter and LuaJIT 2.0 is about factor 15 in geometric mean over all benchmarks, see http://luajit.org/performance_x86.html http://luajit.org/performance_x86.html).
- alberth 6y agoRe: “it seems it only support x86 and x86_64” Probably because it uses DynASM and I believe only those platforms are supported. https://luajit.org/dynasm.html https://luajit.org/dynasm.html
- Rochus 6y ago> Probably because it uses DynASM I had a look at https://github.com/asmjit/asmjit/tree/master/src/asmjit https://github.com/asmjit/asmjit/tree/master/src/asmjit and didn't find any indication that DynASM is used. Anyway, DynASM and LuaJIT are available for many different architectures, not only x86 and x86_64.
- alberth 6y agoThis link has more detail https://github.com/erlang/otp/pull/2745#issuecomment-691482132 https://github.com/erlang/otp/pull/2745#issuecomment-6914821... “ LLVM is much slower at generating code when compared to asmjit. LLVM can do a lot more, but it's main purpose is not to be a JIT compiler. With asmjit we get full control over all the register allocation and can do a lot of simplifications when generating code. On the downside we don't get any of LLVMs built-in optimizations. We also considered using dynasm, but found the tooling that comes with asmjit to be better. “
- moonchild 6y agoRecommend changing the title to '...Erlang's JIT' or similar (though perhaps the 'erlang.org' domain provides enough context?)
- haberman 6y ago> Data may only be kept (passed) in BEAM registers between instructions. > This may seem silly, aren’t machine registers faster? > Yes, but in practice not by much and it would make things more complicated. My understanding is that the latency of an L1 cache load is 5 cycles on recent Intel processors. In tight code that is a really substantial overhead. I get that register allocation is complicated, and I totally see the benefit of simplicity, but it's hard not to think that this will significantly limit performance.
- didibus 6y agoIt does limit performance, but not in a relevant way for the type of applications Beam is designed for.
- CJefferson 6y agoOften, if your function is that small, inlining it is a better idea anyway?
- mathw 6y agoDon't lose sight of two things: Erlang is designed for massive, distributed, highly reliable network systems. Therefore, it is likely that a lot of Erlang code does a lot of waiting around for the network rather than needing every available CPU cycle. This is the first time they're releasing a JIT. They'll deliver their performance gains, and if they want to and it seems feasible they can come back and do further work in the future. But you have to put work in where it's valuable, and somehow I doubt RabbitMQ (for example) is going to care about this being missed this time around.
- a1369209993 6y agoIf the address of the BEAM register file (rbx, apparently) stays constant, and the last write to that register was >5 cycles ago, the load can be speculated 5 cycles in advance; if the last write was <5 cycles ago, my understanding is that that gets handled by store-to-load forwarding with a latency of ~0.
- CalChris 6y agoSkylake Client is "4 cycles for fastest load-to-use (simple pointer accesses) 5 cycles for complex addresses" https://en.wikichip.org/wiki/intel/microarchitectures/skylake_(client) https://en.wikichip.org/wiki/intel/microarchitectures/skylak...
- crazypython 6y agoHow is this a JIT? Everything is compiled to machine code. There's no advanced type specialization or lazy deoptimization. It might as well be an in-memory, lazy, AOT compiler.
- rurban 6y agoA simple method jit is also a good jit, even without escape analysis and advanced optimizations. It avoids the op branching and enables code caching.
- yxhuvud 6y ago> It might as well be an in-memory, lazy Yeah, that is a jit. Perhaps not a very fancy one, the same way ARC is a very unfancy GC, but it is still a jit.
- chrisseaton 6y ago> Everything is compiled to machine code. Yes... but it does that just-in-time, which is what JIT stands for. > There's no advanced type specialization or lazy deoptimization. I don't think those things are required for it to be a JIT are they? I'd say it's a static JIT, rather than a dynamic, profiling, and specialising JIT, but both are JITs. > It might as well be an in-memory, lazy, AOT compiler. Well yeah... that describes a JIT.
- waynesonfire 6y agoreally enjoying this series. excellent job explaining what would be a complicated topic for newcomers.
- PeCaN 6y agoI sort of wonder if this approach to JITing is worth it over just writing a faster interpreter. This is basically like what V8's baseline JIT used to be and they switched to an interpreter without that much of a performance hit (and there's still a lot of potential optimizations for their interpreter). LuaJIT 1's compiler was similar, although somewhat more elaborate, and yet still routinely beaten by LuaJIT 2's interpreter (to be fair LuaJIT 2's interpreter is an insane feat of engineering).
- htgb 6y agoThey detail why they chose the current path rather than continuing improvements of the interpreter in the previous post [1]. See the last few paragraphs for a summary. [1] http://blog.erlang.org/a-closer-look-at-the-interpreter/ http://blog.erlang.org/a-closer-look-at-the-interpreter/