3 ms·
> packaging a library as an npm module using wasm-pack (...) the code would be optimized twice before being executed, first by LLVM and then by the Turbofan JIT
by adrian17 5y ago
> packaging a library as an npm module using wasm-pack (...) the code would be optimized twice before being executed, first by LLVM and then by the Turbofan JIT compiler
Actually, that's not entirely true; if running wasm-pack, there's another optimizer `wasm-opt`, running on output of llvm and wasm-bindgen.
I've found it able to produce measurable improvements in speed and size of the produced .wasm - which is quite surprising to me, given it's basically a set of simple (compared to llvm) optimization passes running on top of whatever llvm already produced, with none of the metadata. However, this also means it can easily do unexpected things like inlining a `#[inline(never)]` function (because again, it only sees the .wasm file, with none of rust or llvm metadata) - I stumbled upon this several times when trying to profile wasm code in browsers.
This also means that if llvm ever supported constant-time semantics, wasm-opt would happily ignore these too.
- codeflo 5y agoYou really do need to look at the whole chain. Of course, if Wasm had constant-time annotations as suggested by the article author, you’d expect wasm-opt to respect those. Then there are steps below that layer to consider as well. Maybe not so relevant when targeting Wasm, but x86 code is increasingly executed inside emulators on ARM chips. Do those respect the constant-timeness? And what about processor microcode and fused instructions, can a future CPU include an optimization that will un-constant-time your code? (In other words: I really don’t know, are there constant-time guarantees made by the ISA?)
- woodruffw 5y ago> (In other words: I really don’t know, are there constant-time guarantees made by the ISA?) Not on AMD64, at least: neither Intel nor AMD will guarantee the timing behavior or timing complexity of an instruction between processors. REP prefixes with string/data operations exemplify this -- they historically had data-dependent timings, but have become increasingly decoupled as both Intel and AMD have shoved "fast string" modes into ucode.