8 ms·
It looks like the slowdown on wasm is related to the powi call in the distance function (which the benchmark page mentions as a likely suspect). Emscripten curr
by sunfish 10y ago
It looks like the slowdown on wasm is related to the powi call in the distance function (which the benchmark page mentions as a likely suspect). Emscripten currently compiles @llvm.powi into a JS Math.pow call, however this powi call is just being used to multiply a number by itself. I filed this issue:
https://github.com/kripken/emscripten-fastcomp/issues/171
to track this issue in Emscripten.
In JS, the assumption is that code may be written by humans, so engines are expected to optimize things like Math.pow calls with small integer exponents implicitly, which is likely why this code is faster in JS. And in the asm.js case, the page mentions that it's using "almost asm", which is not actually asm.js, so it's using the JS optimizations.
This is one of the characteristic differences between JS and WebAssembly: in WebAssembly, the compiler producing the code has a much bigger optimization role to play.
- shurcooL 10y agoFor some reason, Math.pow implementation in JavaScript is really optimized for special cases of multiplying numbers. I ran into this too when I found this surprising case of faster performance of Go when transpired into JS [1]. [1] https://medium.com/gopherjs/surprises-in-gopherjs-performance-4a0a49b04ecd https://medium.com/gopherjs/surprises-in-gopherjs-performanc...
- pcwalton 10y ago> For some reason, Math.pow implementation in JavaScript is really optimized for special cases of multiplying numbers. It's used for multiplying integers in SunSpider: https://github.com/adobe/chromium/blob/master/chrome/test/data/dromaeo/tests/sunspider-math-partial-sums.html https://github.com/adobe/chromium/blob/master/chrome/test/da...
- Twirrim 10y agoOh yay. Micro-optimisations for benchmarks. That could never go wrong.
- kibwen 10y agoBlame the people who continue to hold up SunSpider as a useful benchmark. :P From https://news.ycombinator.com/item?id=7837433 https://news.ycombinator.com/item?id=7837433 : "(For example, did you know that every major JS engine now has a daylight savings offset cache, something which is entirely useless for any real code, but substantially speeds up the date benchmarks in SunSpider? Bleh.)" From https://blog.mozilla.org/javascript/2013/08/01/staring-at-the-sun-dalvik-vs-spidermonkey/ https://blog.mozilla.org/javascript/2013/08/01/staring-at-th... : "All Javascript engines, including SpiderMonkey, use optimizations like transcendental math caches (a cache for operations such as sine, cosine, tangent, etc.) to improve their SunSpider scores."
- kbenson 10y ago> Blame the people who continue to hold up SunSpider as a useful benchmark. As opposed to some other benchmark, that gets some other micro-optimizations? I'm not sure that's productive. That said... > "(For example, did you know that every major JS engine now has a daylight savings offset cache, something which is entirely useless for any real code, but substantially speeds up the date benchmarks in SunSpider? Bleh.)" I'm not sure this is without benefit for me. I store time stamps in UTC in the DB, and pass them along in that form to the client, which then converts to local time. I imagine this might benefit the sometimes hundreds of conversions the client does in my case? Edit: Maybe what we need a something like the Fortune 500, but for webpages. Take the top 50 web properties(by whatever metric), and allow them to submit benchmark webpages that approximate a real world situation. Facebook would have one that loads a canned sample Facebook account, Google might have one to approximate search results, one to approximate docs, and one to approximate gmail, etc. Only allow new submissions once a month or quarter, but index them so we can see performance over time, performance against older technology, and how technologies change over time.
- kibwen 10y ago> As opposed to some other benchmark, that gets some other micro-optimizations? Yes, because SunSpider in particularly is notoriously, egregiously bad. Google's Octane benchmark suite ( https://developers.google.com/octane/benchmark https://developers.google.com/octane/benchmark ) at least attempts to approximate real-world usage by benchmarking actual applications (like PDF.js), as well as scraping regexes from popular sites and using those as part of a regex performance suite. (And even still there are some arguably dubious entries in Octane.)
- galacticpony 10y agoThe whole code is horribly over-engineered, at least for a performance benchmark. Anybody who uses a pow call to square a number must believe compilers are magic. Worse yet, some compilers might actually optimize for that case, reinforcing that belief. Compilers should be dumb, fast and predictable instead. That way, everyone gets what they deserve.
- pcwalton 10y ago> Compilers should be dumb, fast and predictable instead. That way, everyone gets what they deserve. That's a good way to get pummeled in benchmarks. Benchmark scores matter a lot to perception: just look at any HN thread about programming language or browser performance, or any review of browsers in the tech press. We may not like it, but the benchmark-dominated world is the world we live in.
- sqeaky 10y agoI disagree that compilers should be dumb. A smart compiler can do impressive things like take types into account to add optimizations or it can also take use patterns into account to make Javascript nearly as fast as C++. With dumb compilers we are putting the need to understand every machine on every developer. Isolating what knowledge is required in different locations is exactly why we have abstractions. The compiler is just one more abstraction to help this kind of specialization and basic division of labor.
- galacticpony 10y agoThe first paragraph is just nonsense. You, too, suffer from the belief in the magic compiler. Regarding the second statement: These "smart" optimizations happen before architecture specialization, so that's wrong too. But even if it was true, you now shift the burden to understanding every compiler in order to get the optimization you need. If people expect that pow(x,2) transforms to x * x, then every compiler would have to implement that. This is a trivial example, but you in reality you always have to structure your code appropriately for some optimization to kick in, what if one of your target compilers needs a different structure?