12 ms·
Introducing the B3 JIT compiler
- munificent 11y agoReally cool article. Posts like this always make me wonder what the state of the programming would be if browsers hadn't sucked up almost all of the world's compiler optimizers.
- ajross 11y agoTo be fair, GPU vendors sucked up a ton too. But considering that optimized scalar code performance has moved, what, maybe 40% over the last two decades, I'm going to say "not much". Compilers are sexy, but they're very much a solved problem. If we were all forced to get by with the optimized performance we saw from GCC 2.7.2, I think we'd all survive. Most of us wouldn't even notice the change.
- munificent 11y ago> Compilers are sexy, but they're very much a solved problem. Not for all of the other widely-used languages that still have incredibly simple interpreters. Think how much energy could have been saved if Ruby, Python, and PHP were all as fast as your average JS engine.
- VeejayRampay 11y agoExactly. Ruby would probably have massive adoption with JS-like speed.
- ajross 11y agoRuby has massive adoption, as do Python and Perl and PHP. Environments with objectively better performance like the JVM and .NET have not, in fact, done all that well comparatively in this environment (which is to say they've done fine too and achieved "massive" adoption, just not that much better than their slower competitors). In fact, looking at the market as it stands right now I'd say that performance concerns are almost entirely uncorrelated with programming environment adoption.
- yxhuvud 11y agoConsidering the massive speed improvements we have been seeing in Javascript, but also languages like Ruby, I'd say compilers is a solved problem for staticly compiled languages, but perhaps not as much for interpreted highly expressive languages.
- abecedarius 11y agoSort of. Lisp had a good combo of expressiveness and speed in the 80s, but it reached that point on the efficiency frontier by different techniques, like heavier use of macros. The newer techniques like trace compilation can make life even better, but the language design decisions in Ruby/Python/etc. that made that sort of thing necessary if you want speed, they didn't really pay for themselves from the perspective of smug Lisp weenies like me who were happy enough with our language and just wanted pragmatics like libraries.
- sklogic 11y agoDid you just call Javascript a "highly expressive" language? Really?!?
- pizlonator 11y agoLOL! It is expressive though. It's got really powerful first-class functions. It's got classes. It's got prototypes. It's got generators. It's got other things that I don't even remember (but will probably have to learn, to implement them, make them fast, and then fix the bugs). I actually think that the reason why JS is so odd is that it is so expressive. That tends to happen with kitchen sink languages like C++. Relevant: "There are only two kinds of programming languages: those people always bitch about and those nobody uses."
- sklogic 11y agoMore expressive than a less imaginative subset of C. More expressive than assembler. Much less expressive than any decent high-level language. Javascript is low level, primitive and clumsy. There is absolutely no excuse for it being so crappy. If you still think it is "expressive", you never seen an expressive language.
- pizlonator 11y agoajross is right. People choose the languages they like regardless of performance. JS perf is important to users though. Faster execution means fewer watts spent rendering and interacting with your favorite web page. (Fun fact: B3's backend contains a machine description language that gets compiled to C++ code by a ruby script, opcode_generator.rb. We use Ruby a lot.)
- cosinusoidally 11y agoWhy didn't you use JavaScript for that purpose (in a similar vein to how LuaJIT uses Lua for dynasm)?
- pizlonator 11y agoI'm not a big fan of self-hosting. I like that you can build JavaScriptCore without using JavaScriptCore.
- cosinusoidally 11y agoI wasn't really referring to self hosting. I was wondering whether it made sense to write tools like JavaScriptCore's offlineasm in JavaScript rather than Ruby (as LuaJIT does by using Lua for its dynasm tool). The LuaJIT build process actually builds a cut down copy for Lua for the purposes of running dynasm, which is then used to build parts of LuaJIT.
- pizlonator 11y agoWe spend time optimizing JavaScriptCore's build time. It's a hard problem and it's important when managing such a large project. People inevitably have to do clean builds and that sucks when you have >400k lines of C++ code. So, naturally, we want to avoid building JSC twice. ;-)
- iso-8859-1 11y agoSuppose that you are developing not to push this platform, but to simple do better on this platform. What does self-hosting gain you in that case?
- chrisseaton 11y agoMy implementation of Ruby, JRuby+Truffle, is as fast as V8 http://stefan-marr.de/downloads/crystal.html http://stefan-marr.de/downloads/crystal.html
- pizlonator 11y agoNote that usually being "as fast as" a production JSVM means also proving that you can start up as fast as JSVMs do. Have you done this?
- ehsanu1 11y agoSearch around for "Substrate VM". I see it referenced in some slide decks, and it's designed to make JVM startup much faster. Here's an old slidedeck that talks about it: http://www.oracle.com/technetwork/java/jvmls2013wimmer-2014084.pdf http://www.oracle.com/technetwork/java/jvmls2013wimmer-20140...
- ksec 11y agoWow thx for the reminder, I have forgotten about it already. I remember the promise of Truffle + Graal + Substrate VM. My god can't believe it was 3 years ago i read about it on HN.
- pizlonator 11y agoI know about it. I'm not concerned with future hypotheticals, but current actual measurements. That's the currency I deal in.
- ehsanu1 11y agoThere seem to be measurements in the slides I reference, so it wasn't exactly "future work". But of course it sucks that this work is not yet publicly available (AFAIK), probably due to not being 100% production ready yet. But I'm sure it's getting there.
- geodel 11y agoLately quite a few apps are saving energy by moving from Ruby/Python/PHP to Go.
- pcwalton 11y agoAnd Go proves munificent's point: it doesn't have many compiler optimizations either. (This may change with the WIP SSA backend, but the point remains that Go gained huge popularity in spite of having a non-optimizing compiler.)
- pjmlp 11y agoYeah, more interesting is having them rediscovering Turbo Pascal compile speeds. EDIT: I wonder why the positive effect to re-discovering that not all compilers need to be like C and C++ compile speeds and that it was once upon a time mainstream, is worthy of downvotes.
- nickpsecurity 11y agoCan't critique this one: Go was an attempt to re-create the Oberon experience in modern setting with some additions from other languages. Rather than accidental re-discovery, getting Oberon (not Pascal) speed out of the compiler was an explicit design goal. One of few examples of modern IT really learning from the past. Unfortunately, they didn't learn about the stuff between Oberon and 2007 that would've been nice to have in a modern, app language. ;)
- pjmlp 11y agoIt is not a critic, apparently it was understood as such. The remark is tailored to those that think compiled languages can only be slow as C and C++ compilers, since they never used anything else, and then jump of joy when they use Go. Yet if it wasn't for the VM detour of the last 20 years, that experience would probably be a current one, instead of being re-discovered.
- 11y ago
- seanwilson 11y ago> Think how much energy could have been saved if Ruby, Python, and PHP I would think we'd see even better improvements if developers would move towards statically typed languages as well.
- sklogic 11y agoIf Ruby, Python and PHP would suddenly disappear, even more effort and energy would have been saved.
- pcwalton 11y ago> Most of us wouldn't even notice the change. A 40% decrease in optimization is enough to drop framerates from 60fps to 30fps easily, so I'm pretty sure we would notice it.
- mafribe 11y agoCompilers are sexy, but they're very much a solved problem. This may be true for sequential languages, but is very much false for the compilation of concurrency and parallelism. It's basically not known how to do this well. Part of the problem is that CPU architectures with parallel features have not yet stabilised. For sequential languages the problem has shifted: how can I get a performant compiler easily. The most interesting answer to this question is PyPy's meta-tracing, and that's work is from 2009, and far from played out.
- kannanvijayan 11y agoI'd disagree. Classical compiler work is very mature, yes - and new progress in things like register allocation and backend IR-based optimization stuff is well trod ground. But in the context of JIT compilers for dynamically typed languages, in particular the space involving runtime inferred hidden type models, there is a TON of work left on the table. It hasn't been paid much attention to in academia, IMHO largley because of a historical perspective on optimizing dynamic languages as "not classy" work among language theorists. I hope that perspective changes over time.
- nly 11y ago> optimized scalar code performance has moved, what, maybe 40% over the last two decades I'm not convinced. Raw single-thread number crunching performance is somewhere around _two to three fold_, clock-for-clock, on Intel x86, over that of 10-15 years ago. What methodology do you use to attribute only a fraction of those gains to language optimizers? And even if you are correct, why is it meaningful? Who is going to have invested energy in optimising the shit out of mundane codegen when hardware performance will have just come and stolen your thunder a few months later? The problem we have now is that CPUs are gaining ever more complex behaviour, peculiarities, and sensitivities. I'd say compiler engineering is far from a "solved problem", even for statically-typed languages.
- rayiner 11y ago> The problem we have now is that CPUs are gaining ever more complex behaviour, peculiarities, and sensitivities. With mainstream CPUs, exactly the opposite is happening. CPUs are getting more complex under the hood, but less sensitive to code quality. For example, a lot of the scheduling hazards in the P6 microarchitecture have been eliminated in subsequent iterations. Branch delay slots are a thing of the distant past, so are pipeline bubbles for taken branches, indirect branch prediction is extremely capable, even the penalty on unaligned accesses is minimal.
- pcwalton 11y agoWell, sure, but SIMD more than compensates for all of that, given how hard autovectorization is. In fact, I think with things like AVX and NEON becoming ubiquitous, you can get more benefit out of writing in assembly (or intrinsics) than any time I can think of in the past 10 years.
- deleted 11y ago[deleted]
- Joky 11y ago> Compilers are sexy, but they're very much a solved problem No they're not, and won't be for long (ever?). However it does not matter because they are "good enough". Compilers are driven by heuristics which provide "reasonable" results in most cases for common architectures. But they still leave a lot on the table. Compiler writers have to trade compile-time with execution-time. Now we're not talking about an order of magnitude, but rather ~20%-30% in some workloads. When it matters (I guess for people like Google/Facebook/Amazon/... it translates in electricity bill and a number of racks to add to the datacenter) people may have to get down to the assembly level for a very small (and hot) part of the program.
- yxhuvud 11y agoWhile they have sucked up a lot, it is far from certain all would have been employed optimizing compilers if that hadn't happened. Demand tend to create the supply. Also, many of the improvements done for JS will certainly trickle down to Python, Ruby, PHP eventually.
- sklogic 11y agoThey're solving a non-issue. The rest of the world is perfectly fine with the statically typed, well designed languages that are easy to compile. And only the web world is so obssessed with smart compilers compensating (impressively, but still far from being sufficient) for multiple deficiencies in the language design.
- pcwalton 11y ago> The rest of the world is perfectly fine with the statically typed, well designed languages that are easy to compile. Python, Ruby, PHP, and Perl aren't "the rest of the world"? As far as compilers are concerned, all of those languages have more troublesome semantics than JavaScript does. > And only the web world is so obssessed with smart compilers compensating (impressively, but still far from being sufficient) for multiple deficiencies in the language design. You have no idea how much compilers have to compensate for the deficiencies in C and C++'s design.
- sklogic 11y agoNobody really cares about their performance. They're just fine with their simple interpreters. Web is different, there is no choice, no fallback to C. And no, thank you kind sir, but I've got a very good idea of what compilers are doing wrt. C deficiencies, I was writing OpenCL compilers for 6 years at least. Besides aliasing stupidity and byte-addressing there is nothing really bad to compensate for.
- pcwalton 11y ago> Web is different, there is no choice, no fallback to C. asm.js and Web Assembly. > Besides aliasing stupidity and byte-addressing there is nothing really bad to compensate for. Aliasing issues, C++ heavy reliance on virtual methods, too many levels of indirection in the STL, overuse of signed integers due to "int" being easier to type interfering with loop analysis, const being useless for optimization, slow parsing of header files...
- 11y ago
- mrspeaker 11y agoThat seems a bit chicken-and-the-egg-y though: if the web didn't become a global phenomenon then there would be far less interest in improving the browsers, far less business need for programmers, and far fewer people working on whatever they'd be working on if they weren't working on browsers.
- beagle3 11y agoLook at what Mike Pall did with LuaJIT2 - I assume if the world wasn't so focused on the web, we would see more of it in other languages. But really, things aren't that bad. Microsoft has enough good people working on RyuJIT, PyPy has some of the best people advancing metatracing JITs, and Mike Pall is a god among men.
- eyan 11y agoThe Father, The Son, The Holy Ghost, and Mike Pall.
- cpr 11y agoGood to see the Webkit team (mostly Apple) continue putting serious energy into JS performance. Take a bit of guts to throw out the whole LLVM layer in order to get compilation performance... It's also encouraging to see them opening up about future directions rather than just popping full-blown features from the head of Zeus every so often. (Not that they owe us anything... ;-) (Edit: it's also damned impressive for 2 people in 3-4 months.)
- MrBuddyCasino 11y agoGutsy move indeed. Though I wonder, what was the development cost in real life cash for gaining those 5% of performance?
- pizlonator 11y ago3 months. Two people working on it (me and @awfulben).
- titzer 11y agoNice work, Fil. Looks cool.
- pizlonator 11y agoThanks! :-) Looking forward to more write-ups about TurboFan!
- MrBuddyCasino 11y agoOk, thats impressive. I'll show myself out.
- ignoramous 11y agoGreat work! Time to update... https://en.wikipedia.org/wiki/WebKit#JavaScriptCore https://en.wikipedia.org/wiki/WebKit#JavaScriptCore
- bluejekyll 11y ago
- Ecco 11y agoI'm really wondering about the politics behind all this. I mean both LLVM and WebKit are Apple projects (even though they're Open Source). So it would have been reasonable to expect an improvement of LLVM instead of ditching it altogether.
- pcwalton 11y agoI highly doubt there are any politics behind it. LLVM is an AOT C/C++ compiler at its core, and the tradeoffs it makes don't always make sense for dynamic, JIT compiled languages with extreme emphasis on compilation speed like JS. Personally, I expected this to happen.
- zepto 11y agoIf you read the technical explanation you will see that there are zero politics behind this.
- gsg 11y agoSpecialising LLVM for dynamic compilation might well come at the expense of LLVM's strengths as an AoT compiler, and it would probably not be easy in any case. In addition to that, dedicated implementations can take various shortcuts to make their job easier - there are some nice examples given in the link. LuaJIT is example of a compiler project that benefits from being heavily specialised to a particular job, to remarkable effect.
- legulere 11y agotl;dr: B3 will replace LLVM in the FTL JIT of webkit. LLVM isn't performing fast enough for JIT mainly because it's so memory hungry and misses optimisations that depend on javascript semantics. They got an around 5x compile time reduction and from 0% up to around 10% performance boost in general.
- zepto 11y agoActually the bigger reason is compile time - better optimizations based on JavaScript semantics are a secondary advantage.
- pizlonator 11y agoI think that's mostly accurate, in the sense that we wouldn't have done this if it was only motivated by specializing for JavaScript semantics. We had gotten pretty good at having our high-level compiler (DFG) burn away the JavaScript crazy and leave behind fairly tight code for LLVM to optimize. But as soon as we realized that we had such a huge compile time opportunity, of course we optimized the heck out of the new compiler for the kinds of things that we always wished LLVM could do - like very lightweight patchpoints and some opcodes that are an obvious nod for what dynamic languages want (ChillDiv, ChillMod, CheckAdd, CheckSub, CheckMul, etc).
- dochtman 11y agoBut isn't it true that some of the things you ended up doing would make sense for LLVM, or would most of them be invalidated by the kinds of optimization passes that are common in LLVM? E.g. stuff like making the in-memory IR representation better cacheable certainly sounds like it's all-upside, and LLVM should just learn from your project.
- pizlonator 11y agoLLVM's use of a very rich (and hence not as memory efficient) IR is deeply rooted. Phases assume that given any value, you can trace your way to its uses, users, and owners. The LLVM code I've played with assumes this all over the place, so removing the use lists and owner links as B3 does would be super hard. B3 can do it because we started off that way.
- alberth 11y ago>> "tl;dr: B3 will replace LLVM in the FTL JIT of webkit. LLVM isn't performing fast enough for JIT mainly because it's so memory hungry and misses optimisations that depend on javascript semantics. They got an around 5x compile time reduction and from 0% up to around 10% performance boost in general." [1] Is this a knock on LLVM then? I wonder then specifically if this brings to light any concerns over Swift (another dynamic language, and was created by the same person who created LLVM as well). [2] Seems weird that the original creator of LLVM was able to make a dynamic language such as Swift - without any problems. [1] https://news.ycombinator.com/item?id=11105231 https://news.ycombinator.com/item?id=11105231 [2] http://nondot.org/sabre/ http://nondot.org/sabre/
- om2 11y agoSwift is generally ahead-of-time compiled, so compile speed doesn't matter as much. Many of the core parts of Swift also avoid highly dynamic semantics. The most dynamic parts are when interfacing with Objective-C classes, an area that has not been heavily optimized yet.
- pizlonator 11y agoSwift is not a dynamic language. It's statically typed.
- klodolph 11y agoThis comment, and the parent comment, seem to be conflating "dynamic language" with "dynamically-typed language".
- derefr 11y agoWhat's a "dynamic language"?
- barkingcat 11y agoDynamic languages are (generally) not compiled ahead of time. It is possible to have static typing in a dynamic language in the sense that the code is interpreted at runtime (but the types are set statically in the code), and there is no "binary" that one runs. It used to be called interpreted language or "scripting" language - but I think the vocabulary shifted so the word "dynamic" more encompasses what the languages are about.
- ck2 11y agotl;dr https://webkit.org/blog-files/kraken.png https://webkit.org/blog-files/kraken.png https://webkit.org/blog-files/octane.png https://webkit.org/blog-files/octane.png seriously though, dang, how many years of coding to get to that level of expertise
- panic 11y agoCool stuff! Does anyone know why the geometric mean is used for averaging benchmark scores rather than the usual arithmetic mean?
- jsnell 11y agoI would say that geometric mean is the usual way of averaging benchmark scores. It has the property that a given relative speedup on a component benchmark always has the same effect on the aggregate score. With an arithmetic mean the component benchmarks with a longer runtime will dominate the aggregate. Normalizing the results before applying the arithmetic mean doesn't really help either -- the first X% improvement to a component benchmark would still be valued more than the second X% speedup.
- DannyBee 11y agoInterestingly, much of their complaints around pointer chasing, etc, are things LLVM plans on solving in the next 6-8 months. i'm a bit surprised they never bothered to email the mailing list and say "hey guys, any plans to resolve this" before going and doing all of this work. But building new JITs is fun and shiny, so ...
- Alphasite_ 11y agoI imagine they did, considering the head of LLVM is a (probably distant) coworker of theirs.
- DannyBee 11y ago"The head of LLVM" - no such thing exists, but okay. (People have such weird ideas about how these projects work in practice).
- Alphasite_ 11y agoI realise that, but you understand the intent of the comment which is the point.
- Joky 11y agoLLVM instruction selection is slow, there is a "fast-path" which hasn't received much attention (it is only used for -O0 in clang). The new instruction selector work just started and will take a couple of years, considering the tradeoff between spending 3 months on it and waiting a few years for LLVM to be improved (without any guarantee of LLVM reaching the same speed as what they did). See also some thoughts from a LLVM developer on optimizing for high-level languages: http://lists.llvm.org/pipermail/llvm-dev/2016-February/095463.html http://lists.llvm.org/pipermail/llvm-dev/2016-February/09546...
- DannyBee 11y agoThe problem with the path they've taken is it has a finite end. This is why, when they compare it to v8/etc, it's kind of funny. They all have the same curve. Basically all of these things, all of them, end up with roughly the same deficiencies once you cherry pick the low hanging fruit[1], and then they stall out, and get replaced a few years later when someone decides thing X can't do the job, and they need to write a new one. None of them ever get to a truly good state. Rinse, wash, repeat. The only thing these things make real progress, is by doing what LLVM did - someone works on it for years. Let me quote a former colleague at IBM - "there is no secret silver bullet to really good compilers, it's just a lot of long hard work". If you keep resetting that hard work every couple years, that seems ... silly. TL;DR If you really believe they've totally gotten everywhere they need to be in 3 months, i've got a bridge to sell you [1] For example, good loop vectorization and SLP vectorization is hard.
- jlebar 11y agoIt's worth noticing that most of the optimizations here are for "space" -- reducing the working set size or the number of memory accesses. CPUs have gotten much faster than memory blah blah. This is the sort of thing where microbenchmarks may mislead you, because you WSS is probably not realistic. I think we don't have great tools for helping with this sort of optimization. One can use perf to find cache misses, but that doesn't necessarily tell the whole story, as you might blame some other piece of code for causing a miss. Maybe I should try cachegrind again...