9 ms·
A New Bytecode Format for JavaScriptCore
- _alastair 7y agoThis is interesting! If anyone from WebKit is in the comments, can you provide link(s) to the bugs discussing the bytecode caching API? I'd love to take a look but my Bugzilla search skills are evidently weak.
- threeseed 7y agoNot much discussion in the bug report itself: https://bugs.webkit.org/show_bug.cgi?id=187373 https://bugs.webkit.org/show_bug.cgi?id=187373
- tadeuzagallo 7y agoThere are a few bugs related to caching, but the main ones are: The initial implementation of the underlying infrastructure: https://bugs.webkit.org/show_bug.cgi?id=192782 https://bugs.webkit.org/show_bug.cgi?id=192782 Initial C++ and Obj-C APIs: https://bugs.webkit.org/show_bug.cgi?id=193401 https://bugs.webkit.org/show_bug.cgi?id=193401 WIP bug to integrate the cache with WebKit: https://bugs.webkit.org/show_bug.cgi?id=194047 https://bugs.webkit.org/show_bug.cgi?id=194047
- the_duke 7y agoImpressive improvements. I wonder how performance and memory use compare nowadays between V8, Spidermonkey and JavascriptCore. Does anyone have a link to recent trustworthy and thorough benchmarks? (Google doesn't really spit out anything noteworthy...)
- pizlonator 7y agoI believe that JSC is in the lead when it comes to throughput and latency. But that is at least partly based on benchmarks that I had a hand in designing. I don’t know how the VMs stand against each other on memory. JSC is improving in this area a lot recently but I don’t know if it’s just catching up or leaping ahead or whatever.
- brentonator 7y agoIt's first-party but I've found the benchmark results frequently posted in their blog do indeed transform into real-world results. https://v8.dev/blog https://v8.dev/blog
- nielsbot 7y agoWhen they talked about getting rid of their threaded interpreter, I was reminded of this article from 2008 about writing fast interpreters, if that's interesting to this audience: https://news.ycombinator.com/item?id=2593095 https://news.ycombinator.com/item?id=2593095 [hn link]
- cogman10 7y agoI'm unclear, why transform to a bytecode as a first step? Wouldn't it be simpler to instead to transform to an AST and work against that for everything? Wouldn't it make sense to generate the bytecode as sort of a last step before doing heavy duty optimizations? Seems like with JS you would be constantly transforming from bytecode -> AST -> bytecode -> machine code, such as every time a method adds a new field to an object or some optimization assumption is violated. It doesn't seem like it would be easier or faster to work with. I'm guessing there wouldn't be a whole lot of memory benefits either. Would someone mind illuminating me on this design choice? (Granted, I've not taken a compilers course, so feel free to call me an idiot for not knowing something basic about how compilers work).
- steveklabnik 7y agoA bytecode is easier to keep stable than an AST is; an AST changes as the language changes, but that doesn't mean you have to change the bytecode, only change the transformation from the AST to the bytecode. I have no idea if this is the reason, but it's a possible one.
- pizlonator 7y agoThat’s a pretty good reason also.
- mhh__ 7y agoIt's a bit of a micro-optimisation but ASTs are also usually represented fairly sparsely, so processing on bytecode could be much faster due to more efficient cache utilisation etc - some benefits from memory locality and also from bytecode usually being smaller in memory) (I don't know JSC does it so that could be wrong)
- deleted 7y ago[deleted]
- pizlonator 7y agoAST is less transformable. We transform our bytecode as if it was just an IR. ASTs are a pretty annoying form to use as a source of shared truth for OSR. JSC has 4 tiers so we need a convenient-to-use shared truth IR. That’s what bytecode is for. For example, bytecode offers high-scalability answers to questions like “where should I exit” and “what is live when I get there”.
- riotman 7y agoDo they have a different bytecode and runtime from WASM? Why not unify everything to web assembly byte code?
- pizlonator 7y agoThey are different. Unifying is not a bad idea.
- ori_b 7y agoWASM byte code is structured into a tree, which means an interpreter loop wouldn't perform well. You'd really want to flatten it into a simpler bytecode before interpreting it -- and you'd want to do other transforms on it before optimizing it. I don't think WASM bytecode is a good format for execution, and it's only mediocre as a compilation target.
- riotman 7y agoThen why not abandon wasm, and make this byte code a target for llvm?
- ori_b 7y agoLLVM bitcode is CPU and ABI dependent, and isn't even stable between LLVM releases.
- riotman 7y agoBad question on my part. More relevant: Why wasm? From what I gather, js seems to already have a bytecode for their JIT runtime. Why not expose that so that we can have C++ in the web?
- singularity2001 7y agoI asked the same question, see below for interesting thread
- pier25 7y agoIs it possible to pregenerate that bytecode? For example for hybrid desktop/mobile apps.
- eridius 7y agoIf you're going to precompile your JS, why would you want to emit this bytecode instead of just going to WASM?
- pier25 7y agoBecause wasm still doesn't support many features that JS does. For example accessing the DOM.
- firethief 7y agoBytecode is a broad category. The design of a bytecode suitable for interpreting JS has little in common with a bytecode for AOT compilation of static languages without tracing GC, other than that they're both made of bytes.
- olliej 7y agoNo, because that would require exposing implementation details. e.g. the first API version of the bytecode would be the fixed API, and any future changes would require a translation layer for converting from one version of the bytecode to the other. The bytecode is also an internal format so is consider trusted - once you expose it you have to worry about invalid bytecode.
- novok 7y agoYou mean like binast? https://github.com/binast/binjs-ref https://github.com/binast/binjs-ref
- olliej 7y agobin ast is/was an attempt to make a more compact version of JS that was easier to parse. But parsing JS isn't a significant bottleneck, it's the subsequent work to produce the stuff that actually runs. IIRC proponents also claimed it meant you didn't have to worry as much about validation, which isn't true because it's untrusted content that comes from the internet. It must be validated. binast also is not designed to be any more readily executable that JS - so even if it were supported step one would be "produce a bytecode that is fast".
- twoodfin 7y agoRe: Direct vs. indirect threading (aka ordinary dispatch). I had the impression that recent Intel x86 chips did enough trace caching to render any performance distinction basically irrelevant. Is that right? (Of course, Apple has their own ARM implementations to consider.) EDIT: Here’s the paper I was thinking of in this regard: https://hal.inria.fr/hal-01100647/document https://hal.inria.fr/hal-01100647/document
- wahern 7y agoMy first thought when they said they switched from direct threading to indirect threading was that it may have been the worst possible time to do that given Spectre. The negligible performance difference on modern Intel chips between direct and indirect threading is largely a result of their deep, heavily buffered branch predictors. Spectre mitigations are going to increase the performance differences. Most userspace applications will forgo mitigations in favor of keeping performance, but the browser was the big exception as it has to be concerned about in-process side-channels.
- eridius 7y agoHaven't all the browser engines already implemented their own mitigations for Spectre/Meltdown for JavaScript code?
- wahern 7y agoI can't speak to their existing mitigations directly, but more generally Spectre will be an ongoing saga for years to come. I seriously doubt the mitigations in place, whatever they are, are comprehensive even for existing proven Spectre channels. The engines haven't even solved RowHammer. It's a very difficult problem domain, particularly for an application JIT'ing random, untrusted code from the Internet. In many respects they're (IMHO) quietly punting because there just aren't satisfactory solutions. Browsers are really stuck between a rock and a hard place, more so than VM providers like AWS EC2. One of the more general solutions is simply to stop trusting the browser, if not your entire operating system. Keep your most sensitive secrets inaccessible to software to the greatest degree possible. Use smartcards and other hardware tokens for authentication, for example.
- DiseasedBadger 7y agoWebKit2 Qt when? Also, full QNetworkConnection support. That would be great.
- singularity2001 7y agoNow please expose a loadFromByteCode api so that we can target bytecode instead of transpiling to js.
- Conlectus 7y agoThis is, give or take, what web assembly is supposed to do, no?
- singularity2001 7y agoWasm is a bytecode, but not js bytecode. In fact compiling js to wasm is very nontrivial / requires compiling it with a whole VM.
- olliej 7y agoThe bytecode is an implementation detail that can, and does change. Making it an API would necessarily make changes (like this one) impossible.
- singularity2001 7y agosome standardization wouldn't harm. Then we can do loadFromByteCodeV1 and later loadFromByteCodeV2
- olliej 7y agoNo, it absolutely would. Because then any changes require duplicating the byte code parsers, and adding translation logic to manage those changes. You must understand: any change would be an API break, and API cannot change. Deprecating an API takes years. You also can't change/add API in security updates, which means if a security fix required changing some aspect of the bytecode that change couldn't be made. Lets say you change numbering of registers? Thats an API change. Lets say you change number of opcodes? API change. Any new API in macOS and iOS requires weeks of review cycle to ensure: * It solves a problem in a generic manner: e.g. it doesn't solve a very specific version of a general issue * It can be kept stable: e.g. it doesn't expose implementation details that may change * It is ABI stable: direct memory access APIs cannot expose layout that can vary based on any internal details. Shipping API is incredibly difficult if you care about ABI stability for software.
- ksec 7y agoAnd it is already shipped in Safari 12.1 ( Which means many are already using it ) Time and Time again it seems WebKit are the only team that wants to make the Web with better Web Pages with javascript , All the others seems to want the Web to be Fat Apps. In the hope of anyone in Safari team is reading it. Please make the Tab Overview Cache the Thumbnail or in List format, currently pressing Tab Overview will reload all the tabs in the background. I don't know if this is for generating thumbnails or other reason. Something that kills my machine when I have 300 Tabs, most of them are "cold" and not loaded.
- saagarjha 7y ago> Time and Time again it seems WebKit are the only team that wants to make the Web with better Web Pages ( And we are far from perfecting it ), All the others seems to want the Web to be Apps. I don’t see how this blog post supports that view, since it’s talking about optimizing JavaScript.
- ksec 7y agoBy Web page I mean including minimal Javascript, such as making initial JS loading faster, lower latency, lower memory, and in general much better UX. To me Chrome seems to be optimising for the wrong thing, like maximum throughput, WASM, Super fast in Compute intensive usage but eating memory like crazy.
- dchest 7y agohttps://v8.dev/blog/ignition-interpreter https://v8.dev/blog/ignition-interpreter "V8 team has built a new JavaScript interpreter, called Ignition, which can replace V8’s baseline compiler, executing code with less memory overhead and paving the way for a simpler script execution pipeline." https://v8.dev/blog/preparser https://v8.dev/blog/preparser "Lazy parsing speeds up startup and reduces memory overhead of applications that ship more code than they need." https://v8.dev/blog/embedded-builtins https://v8.dev/blog/embedded-builtins "V8 built-in functions (builtins) consume memory in every instance of V8. The builtin count, average size, and the number of V8 instances per Chrome browser tab have been growing significantly. This blog post describes how we reduced the median V8 heap size per website by 19% over the past year." https://v8.dev/blog/improved-code-caching https://v8.dev/blog/improved-code-caching "V8 uses code caching to cache the generated code for frequently-used scripts. Starting with Chrome 66, we are caching more code by generating the cache after top-level execution. This leads to a 20-40% reduction in parse and compilation time during the initial load." Etc, etc.
- akling 7y agoKudos to Tadeu Zagallo and the JSC team for landing this awesome patch! I've worked on WebKit memory performance in the past, so I'm well aware that these aren't low hanging fruits we're talking about. The type safety bonus features look great too. :)