11 ms·
Staring at the Sun: Dalvik vs. Asm.js vs. Native
- JoachimSchipper 13y agoOn one hand, it's great that the author did all this work. On the other hand, the benchmark is quite dubious: aside from the Javascript-engines-are-heavily-optimized-for-SunSpider thingy (which, if you click the link at the end of the article, ends up often making plain JS faster than asm.js), the most obvious port of JS code is likely not the most obvious way to write Java/C++, let alone the most performant way. Still, the fact that you can compare Javascript and C++ without needing a log scale is quite an achievement.
- krallja 13y agoThe asm.js is not implemented by hand; it is compiled from C++ via Emscripten, so there's no "obvious port of JS code" - it's all done by the compiler.
- duaneb 13y agoDid the author write the C++ code? This is all pretty useless without being able to reproduce it. EDIT: I'll eat my hat, the code looks great. EDIT2: Looks like he allocates memory in an inner loop, no wonder native is so slow... I wouldn't take the native benchmarks seriously at all. https://github.com/kannanvijayan/benchdalvik/blob/master/nsieve/nsieve.cc#L38 https://github.com/kannanvijayan/benchdalvik/blob/master/nsi...
- kannanvijayan 13y agoI tried to keep this as faithful as possible to the original JS code I was copying from. In the actual SunSpider benchmark (in Javascript), a new Array object is allocated within that same loop, so I wrote the C++ and Java code to mimic that behaviour. I could have pulled the Array allocation out across all of the implementations, but I tried to avoid making any changes to the benchmarks unless there was a correctness-issue involved (e.g. moving makeCumulative in the fasta benchmark out to the prelude was a correctness issue.. since it's wrong to run it more than once on the same array).
- duaneb 13y agoPerhaps, but that's still something you just wouldn't do in native code. If you were to write code that way you wouldn't be writing C++ in the first place (I would hope). I understand the reasoning, but I also think it's misleading in terms of results.
- Skinney 13y agoWouldn't it be more misleading if the C++ version did something different?
- kannanvijayan 13y agoIt's not how you would write it in Javascript either, or Java. Actually, in general you wouldn't be implementing a sieve algorithm at all. That's a general pitfall of benchmarks like these - the micros don't test real programs so much as they test a limited set of implementation mechanisms. In this case, what we're measuring is: "allocate an array, fill it, and then scan it with various stride lengths and mutate it, and then free it.. what does that cost on average, given this spread of array sizes?" The useful thing with these benchmarks isn't what the final numbers are, but why they are what they are, and what that suggests about the underlying implementation. To put it another way, I think these sorts of comparisons are more useful for being able to confirm that some set of mechanisms work roughly equivalently in one vs the other system.. rather than useful for saying one is "better" than the other in any objective way. For example, the nsieve result suggested an issue with ARM codegeneration with asm.js. But the fact that scores between asm.js and native are pretty close on x86 desktops suggests that outside of codegeneration, the cost of allocating arrays, scanning them, and mutating them like this is roughly equivalent on asm.js and native. Similarly, the nbodies result might be suggesting that double-indirection in hot code is a weak spot for asm.js compared to native. The fasta result suggests that there are high overheads associated with using Java collections for small lookup tables of primitive values. With benchmarks, my opinion is that the scores themselves are less important than how you interpret them. Thus I'm not as concerned about how one would optimally implement a looped sieve algorithm in C++ vs Java vs JS, since that's not what I'm trying to get at.
- chrisaycock 13y agoThe point about memory allocation is similar to Walter Bright's explanation of D compiler performance [1]. There, the issue was deallocation; Walter "cheated" by never deallocating any memory since the compiler is a short-lived process. As a complement in this Mozilla test, Kannan Vijayan believes that Asm.js ran so fast in the Binary Trees test because there is a single large allocation at process startup, much like a memory pool. The moral of the story is always be aware of what malloc() and free() are doing to your code. [1] https://news.ycombinator.com/item?id=6103883 https://news.ycombinator.com/item?id=6103883
- tehwalrus 13y agoI once made a C python extension considerably faster by being more careful with malloc and free - I can recommend it as an easy-ish way to speed up a C/C++ program.
- duaneb 13y agoAnyone who has implemented a memory allocator knows how expensive it can be. If you have a simpler algorithm that works with your data, allocate a huge chunk and manage it yourself. EDIT: Looks like the author of the native benchmarks needs to learn this lesson: https://github.com/kannanvijayan/benchdalvik/blob/master/nsieve/nsieve.cc#L38 https://github.com/kannanvijayan/benchdalvik/blob/master/nsi...
- Peaker 13y agoAnd use so-called "intrusive" data structures and APIs to avoid more allocations.
- deleted 13y ago[deleted]
- vidarh 13y agoMy favourite example of memory allocation anti-patterns: A decade or so ago I worked on a GUI and we needed a Type 1 font parser. Freetype was not a good fit to our (very memory constrained, slow) platform. I was suprised though, how slow the library was, and did some profiling. First thing I noticed was a huge number of malloc() calls. And they were all small. A huge portion of them were only 4 bytes. The malloc() implementation on our platform had an overhead of at least 12 bytes per allocation, so not only was it slow, it also wasted a ludicrous amount of memory. All in all there were multiple allocations per glyph despite the fonts being loaded and unloaded as a single unit. Yikes. Since we were in a rush, my hacky fix was a search/replace for malloc()/free() to special a special version of malloc that would just grab the next free chunk from a pool, or if the pool was full allocate another 4KB block and start allocating from that, and nothing for free, and then I added code to initialise the pool and free it to the open/close parts of the font API. The result was drastically reduced memory usage and at least an order of magnitude speedup on our platform. (The library in question was t1lib, a library originally released by Adobe, and the stupid memory handling is still in there as of today despite sporadic updates over the last decade; I guess it's not used much any more with FreeType, but I'm almost tempted to do a cleaner fix and submit it - the fix we did was not suitable for general usage as it explicitly assumed a single set of fonts would be loaded and unloaded in the same order, as it didn't keep pools per loaded font)
- peterhunt 13y agoI've been watching a bunch of this JS performance stuff unfold over the past few months (including that huge rant a while ago). People lament the lack of JITing in UIWebView or that JS is inherently slow or whatever but if you're targeting mobile, there's usually only one thing that matters: Can your JS rendering code consistently execute within 16ms? JS (even without JIT) is certainly fast enough to do this if you offload anything intensive to workers (in fact, I recommend that you put your whole app except for the real time aspects in a worker if possible) and schedule long running tasks over several requestAnimationFrames. Usually the only issue is that GC pauses can cause hiccups > 16ms and you're in no control of that. This has traditionally been seen as a deal-breaker. That's why I'm excited about Asm.js (and LLJS) -- even on generic JS runtimes, it's my understanding they don't generate garbage, and can execute without GC pauses. So I'm looking forward to writing most of my app in traditional JS in a worker, and the realtime components in Asm.js.
- munificent 13y ago> if you're targeting mobile, there's usually only one thing that matters: There are two things that matter, actually: 1. How quickly can you get your app up and running and responding to user input? 2. Once that's done, can you continue to maintain 60 fps?
- peterhunt 13y agoGood point. I was mostly responding to common criticisms. I think JS boot time is usually pretty good since it's easy to incrementally load code in.
- lukifer 13y agoIt's not just about loading the JS; it's about loading the entire WebView, or whatever other environment. (Obviously, this is an issue that every application faces, it's just something worth paying attention to.)
- modeless 13y agoGC pauses are not the only issue. Image loading often introduces longer pauses than GC, and input latency is a huge problem too. I've written a benchmark that exposes these responsiveness issues: http://google.github.io/latency-benchmark http://google.github.io/latency-benchmark
- tyre 13y agoThis is a rather dubious comparison and seems more targeted at selling ASM.js than making a meaningful comparison. The author includes asm.js but not native JS. Well, he actually does test it, but removes it from the results because it did too well. From the article: To be frank, I didn’t include the regular Javascript scores in the results because regular Javascript did far too well, and I felt that including those scores would actually confuse the analysis instead of help it. https://blog.mozilla.org/javascript/files/2013/08/Dalvik-vs-ASM-vs-Native-vs-JS.png https://blog.mozilla.org/javascript/files/2013/08/Dalvik-vs-...
- kannanvijayan 13y agoAuthor here. The "regular JS scores" do REALLY well on several benchmarks (faster than native, even), for one primary reason: the transcendental math cache (which every JS engine uses). This is a very specific optimization that expects that functions like "sin", "cos", and "tan" will be called repeatedly with the same inputs, and puts a cache in front of those functions. In sunspider, this helps. In the real world, we call this "overoptimization for sunspider". I've said this before, and I'll say it again: Sunspider is a poor benchmark to use to talk about JS engines. All of them game it - ALL of them. If I _had_ included the regular JS sunspider scores, the comparison would be unfair since all the JS engines are specifically optimized for sunspider. The reason these benchmarks are somewhat appropriate for Java and C++ is _because_ Java and C++ compilers and libraries have not been optimized with sunspider in mind, and things like the transcendental math cache don't skew numbers. (Also, to be clear: OdinMonkey, the asm.js compiler in SpiderMonkey, does NOT use the transcendental math cache like the optimizer for regular JS code does. This is why "plain JS" is faster than asm.js in several of the benchmarks).
- mraleph 13y ago> for one primary reason: the transcendental math cache If this is the primary reason then it implies that those benchmarks primarily measure performance of transcendental operations with repetitive arguments and thus nothing interesting. Why do you even bother include them in this case? > Sunspider is a poor benchmark to use to talk about JS engines If you think it is a poor benchmark why do you use it or remind outside world of its existence? In my opinion SunSpider is indeed a poor benchmark in general to talk about any kind of adaptive JIT (which includes JVM). That is why I never use it for anything.
- justinsb 13y agoKudos to the author for including their code. Too many benchmarks don't. And it is amazing that JS is even in the same ballpark as Java. That said, like all benchmarks, there are systematic biases. I looked at the binary trees benchmark. The obvious problem is that it uses far too few iterations (100), so the runtime was 60 milliseconds (on my laptop). That's really not enough time for JIT to kick in, although probably JS does JIT more eagerly than Java. I upped it to 10000 iterations: (OpenJDK) Java took 3.3s, JS (with node) took 7.8s, C++ took 15s. (C++ is really hurt by garbage collection vs alloc/free.) Switching C++ to use Google's TCMalloc brought it down to 10s. When you see a benchmark that says that X is faster than Java, and X does not include the letter C, take it with a pinch of salt!
- janjongboom 13y agoBut now you're comparing unoptimized javascript to java, instead of asm.js.
- icebraining 13y agoAnd he's comparing JavaScript with the JVM, not Dalvik. And on a laptop.
- justinsb 13y agoAh - fair point. How can I repeat the asm.js results on my laptop?
- azakai 13y agoGet the benchmark code, get emscripten, and compile them, emcc -O2 nsieve.cpp for example. Yes, I also had to increase the runtime, they were quite short on a laptop - they were meant for a phone I think.
- nxn 13y agoNot sure if this is relevant these days, but the Emscripten FAQ mentions the need to use "-s ASM_JS=1" as an argument to emcc in order for it to actually output asm.js style code. See "Q. How fast will the compiled code be?" from: https://github.com/kripken/emscripten/wiki/FAQ https://github.com/kripken/emscripten/wiki/FAQ
- justinsb 13y agoThe spectral norm test appears to be wrong. On C++ & Java it gets the math expression in the A function wrong, and produces NaN, which apparently has a huge performance cost. Taking the correct version from the Javascript (and converting to a double correctly) gives the correct result, and (OpenJDK) Java runs twice as fast.
- justinsb 13y agoKannan, if you want to fix it, the Java diff is: - return ((double)1)/((i+j)(i+j+1)/((double)2+i+1)); + return 1.0/((i+j)(i+j+1)/2+i+1); That gives the same count as the JS. It'd be interesting to know whether Dalvik/ARM has the same performance hit on NaN!
- kannanvijayan 13y agoThanks! I think the "/2" should be a "/((double)2)", since in C and Java it'll treat the former as a truncated int divide. I do see the issue with the precedence on the 2+i+1, though. Will fix up and re-run.
- justinsb 13y agoYou're quite right. You can also use 2.0
- kannanvijayan 13y agoFixed up and pushed to repo. If you see anything else that's incorrect, please feel free to let me know either by posting or through mail, and thanks again for pointing that out. With the changes, the scores didn't change dramatically, and the asm.js version got a bit faster too, actually coming in closer to the native score than before, which is weird. Overall, scores got better across the board by about 8% or so. The fact that asm.js does better now than it did before somehow suggests to me that there may be an issue in the NDK libc's malloc or free implementation, and that the better scores are simply allowing that issue to have more effect. This is another one of those programs which does a bunch of "biggish" allocs/frees repeatedly. I'll put the updated charts up with an edit note shortly.
- kin3tic 13y agoI'd like to see Rust thrown in the mix.
- zbowling 13y agoSo we have a better NDK than Googles at Apportable. We use our own malloc and better C++ runtime (libc++ instead of libstdc++) and use Clang 3.4 instead of GCC. One problem with your tests is that you are using the system malloc, and that is horribly slow (and dalvik gc will even obtain a global lock on it every now and then). Firefox does not use the system malloc (instead it uses jemalloc). This actually has a big time savings in tests that will call malloc at any point. I would love to run your tests on our platform. Can publish the exact times you got? I want to spin it up and see if I can get better numbers on the native side.
- kannanvijayan 13y agoSure, here's a paste of a CSV file containing the data I recorded. This was taken on a Nexus 4 running Android 4.2.2. There are two columns for each bench - the right column containing nine individual scores, and the left column containing their aggregate stats: http://paste.ubuntu.com/5960016/ http://paste.ubuntu.com/5960016/
- azakai 13y agoAll the code to the benchmarks is available, and all the tools use to compile them are open source, so please make that comparison, would be interesting to see!
- MBCook 13y agoInteresting. It's nice to see someone actually do analysis instead of just dumping numbers and saying "So X is the fastest, except when it's Y". I'd like to see all these re-run on a desktop with a desktop JVM. I'm curious if the problems Dalvik showed on the binary tree bench are specific Dalvik specific or if Hotspot has optimizations that fix some of the issues identified.
- kenster07 13y agoIt is quite misleading to write native code in the same way you write the javascript. A wothwhile benchmark would test optimized code in all contestant languages.
- kenster07 13y agoWas this point downvoted for a good reason? I will reiterate my point. There is no point in comparing suboptimal C++ to optimized asm. Rather than downvoting me, the author if the study ought to rerun the benchmark with optimized C++.
- lttlrck 13y agoIts a shame the benchmarks chosen did not allow useful comparison between Asm.js and plain Javascript to be included. Plain JS is an important baseline.
- fuckjavascript 13y agoWell, it's good to see that the browser developers have finally conceded total defeat on Javascript as the basis for their platform, and are now simply constructing a nice little low-level virtual machine for general compilation. Granted, they're still hypocritically trying to keep up appearances by sharing as much machinery as possible with their Javascript VM, but nobody's perfect, I guess.
- fuckjavascript 13y agoSnarky tone aside, care to explain what is inaccurate or unhelpful in this comment? There is a lesson to be learned from the debacle of browser design about ignoring what worked in the past. Does anyone seriously dispute, now that they have FINALLY constructed a VM for low-level code execution combined with some primitive APIs, that the browser as a document viewer with Javascript and DOM banged on is a completely and utterly BUNGLED design? Once you have a GOOD programming environment you can simply create your document viewer as an application inside it! And now that it's 2013 I expect more than old ideas done badly. For example, at this point I would go further and say that the sandboxing in the browser should be handled by the VMM, not by the optimizer! The language should provide portability, not safety. That way optimizations can be as aggressive as you like without worrying about security. But hey let's throw a whole document browser and JIT into the trusted base of this distributed computing platform. I got nothing but ridicule and talk about how the "web" was special from people here on HN in the past when I tried to make these suggestions. Turns out this was utter garbage and the web is like any other programming platform - except designed and used by absolute chimps.