15 ms·
If you were ever curious how modern JavaScript VMs (or VMs for other dynamic languages) achieve high performance, this is an awesome resource. It explains tiers
by om2 6y ago
If you were ever curious how modern JavaScript VMs (or VMs for other dynamic languages) achieve high performance, this is an awesome resource. It explains tiers, the goals and design of different tiers, on stack replacement, profiling, speculation and more!
JavaScript engines are the most advanced at this (only LuaJIT is even comparable), it would be awesome if Python, Perl, Ruby, PHP or the like aimed for the same level of performance tech.
- 1337shadow 6y agoActually that's available for the Python language as well with PyPy.
- pizlonator 6y agoThe post gives PyPy a shout out. But it’s subtle. PyPy is similar but not the same. I think that JSC’s exact technique could be tried for Python and I don’t believe it has.
- 1337shadow 6y agoMaybe, I wish I had the occasion to deploy PyPy in production, but as a daily Python user since 2008: I never had to switch to PyPy to fix any performance problem that mattered for my users or me, but I keep an eye on PyPy and admire it as well as the developers behind it.
- hajile 6y agoPython actually uses prototypal inheritance behind the scenes. Pypy is a tracing JIT though which is quite different. JSC and v8 compile a method at a time. Spidermonkey used to use tracing, but switched to methods too (though I think it still does limited tracing in some situations).
- jashmatthews 6y agoTraceMonkey was removed almost a decade ago https://bugzilla.mozilla.org/show_bug.cgi?id=698201 https://bugzilla.mozilla.org/show_bug.cgi?id=698201
- hajile 6y agoJaegerMonkey used both approaches. I don't know if IonMonkey still does though. https://hacks.mozilla.org/2010/03/improving-javascript-performance-with-jagermonkey/ https://hacks.mozilla.org/2010/03/improving-javascript-perfo... EDIT: It looks like a couple trace-like techniques are still used in IonMonkey (maybe the devs have some input there though). https://wiki.mozilla.org/IonMonkey/Overview#Modus_Operandi https://wiki.mozilla.org/IonMonkey/Overview#Modus_Operandi
- jashmatthews 6y agoThere’s no tracing in either JägerMonkey or IonMonkey. That’s just a confusing statement about the similarity of speculation guards and OSR exits to trace guards and side exits. For a while TraceMonkey was used as a higher tier above JägerMonkey which might be what you’re thinking of? I don’t work for Mozilla but spent a ton of time studying TraceMonkey for building my own tracing JIT.
- goatlover 6y agoThere is also Julia, which JITs very performant code.
- slaymaker1907 6y agoI think that Julia uses the LLVM JIT. When I used it last, Julia was somewhat prone to long startup delays since it would fully compile functions on first execution. That can work great for long running applications such as its main niche of scientific computing but would be terrible for JS since you want the page to be interactive ASAP. This is why talking about JIT performance is so complicated. Not only do you need to worry about compilation speed and speed of the generated code, you also have to worry about a lot about impact on memory and impact on concurrently running code. Plus most JITs also need to have some sort of profiling system running all the time as part of those constraints to only spend compilation resources on hot paths.
- pizlonator 6y agoThe difference between Julia and speculative compilers is that Julia is much more likely to have to compile things to run them. JSC compiles only a fraction of what it runs, the rest gets interpreted.
- om2 6y agoModern JS engines have a multi-tier structure and profiling info that lets them choose what to JIT and which point on compile speed vs runtime speed tradeoff space to take for any given chunk of code. The post covers a lot of this.
- tyingq 6y agoPHP has a JIT coming in 8.0, using the same underlying tech that LuaJit does. Unfortunately, most of what people do with PHP isn't CPU bound, so it doesn't help much.
- pizlonator 6y agoAnd let’s not forget HHVM. That was (is?) quite the beast.
- pastrami_panda 6y ago> most of what people do with PHP isn't CPU bound ... cause it's I/O bound? Just curious.
- pizlonator 6y agoI feel like I can almost visualize the killshot slides that the HipHop and HHVM folks used to show the extent to which PHP can be CPU bound and the extent to which server idle time can be increased by using speculative compilation tricks on PHP. So, I think it is cpu bound for at least some people.
- tyingq 6y agoVersus PHP5.x that's true. Mostly because the PHP internals were inefficient. The performance differences vs HHVM vanished with PHP 7.x.
- jashmatthews 6y agoHHVM's 2nd generation region based JIT performs enormously better than PHP7.x in places where dynamic languages don't perform so well but PHP7.x has the clear lead when benchmarking large apps like WordPress.
- tyingq 6y agoI get where the "rooting for the underdog" feeling comes from, but it still feels good that the relatively small and underfunded PHP team mostly beat Facebook here. I like to imagine there is some internal debate at FB on whether to just go back to mainline PHP and kill HHVM. Especially with a credible JIT coming.
- monocasa 6y agoI'd say that the JVMs (particularly Azul's) are more advanced at this. Even with types, they sill speculate in order to inline across virtual method calls. But agreed that the amount of perf JS engines can achieve is truly impressive.
- pizlonator 6y agoJSC speculates in order to inline across virtual method calls while also inferring types. Also inlining across virtual method calls is just old hat. See the post’s related work section to learn some of the history. Most of the post is about techniques that are more involved that inlining and devirtualization.
- monocasa 6y agoHotSpot is a fork of strongtalk, which did the same thing and in fact invented the techniques you're talking about (ie. creating fake backing types for untyped code and optimistically inlining those with bounces out to the interpreter on failures, perhaps keeeping several copies around and swapping out the entry in the inline cache). Additionally that functionality has been added to over time with the invokedynamic byte code and it's optimizations.
- pizlonator 6y agoInvented some of the techniques in the post. Most of the post is about what’s new. I cite the work that led up to strongtalk throughout the post.
- monocasa 6y agoI read the article and don't see any examples of what hotspot doesn't do? What am I missing?
- pizlonator 6y agoHotSpot has four tiers huh? That’s just one snarky example. There are lots of others. I’m sure you could identify them easily, if you are familiar with HotSpot and you read the post.
- jashmatthews 6y agoLuaJIT isn't really comparable. Filip touched on why the tracing compilers struggle: tail duplication. Avoiding tail duplication means having heuristics which keep traces short which means limiting your optimization scope. You can see an example of working around the limited optimization scope by templating Lua here: https://github.com/LuaJIT/LuaJIT-test-cleanup/blob/master/bench/fasta.lua#L21 https://github.com/LuaJIT/LuaJIT-test-cleanup/blob/master/be... This makes some variables which would otherwise have to be loaded at the start of the trace into constants in the recorded trace. A more complex compiler can just do the same optimizations without the fuckaround.
- om2 6y agoMaybe not in architecture and thus in perf ceiling, but they actually get good perf, which is generally not the case for the main implementations of other popular scripting languages IMO. Thanks for highlighting the tracing JIT issue.
- anonymoushn 6y agoIs that really the point of that code? It seems like the point is to generate code to map real numbers from math.random to letters in a probability distribution in the smallest number of branches.