4 ms·
Some of the performance gain was coming from the use of Unsafe in earlier versions of Spark (e.g. Spark 1.5). However, the massive gain you are seeing in Spark
by rxin 10y ago
Some of the performance gain was coming from the use of Unsafe in earlier versions of Spark (e.g. Spark 1.5). However, the massive gain you are seeing in Spark 2.0 are not coming from Unsafe. It is coming from this idea we call "whole-stage code generation", which eliminates virtual function calls and puts intermediate data in CPU registers as much as possible (versus L1/L2/L3 cache or memory).
We will be writing a deep dive blog post about this in the next week or two to talk more about this idea.
- cachemiss 10y agoSeems similar to this: http://www.vldb.org/pvldb/vol4/p539-neumann.pdf http://www.vldb.org/pvldb/vol4/p539-neumann.pdf
- rxin 10y agoYup similar to that. Our next blog post in the pipeline is going to reference this paper.
- cachemiss 10y ago(Guessing you meant to respond to me) Excellent, I'm glad that the "big data" world is starting to look at database literature in terms of how it does execution, as there is much to be learned. Most of these systems are extremely inefficient (looking at you Hadoop), when they don't really have to be. Efficient code generation should be table stakes for any serious processing framework IMO.
- rxin 10y agoYup was replying to you. Clicked the wrong button :)