3 ms·
AFAIK, there are only a handful of DBMSs that do complete query compilation with the LLVM: * MemSQL (http://highscalability.com/blog/2016/9/7/code-generation-t
by assface 9y ago
AFAIK, there are only a handful of DBMSs that do complete query compilation with the LLVM:
* MemSQL (http://highscalability.com/blog/2016/9/7/code-generation-the-inner-sanctum-of-database-performance.html http://highscalability.com/blog/2016/9/7/code-generation-the...)
* Tableau/TUM HyPer (https://blog.acolyer.org/2016/05/23/efficiently-compiling-efficient-query-plans-for-modern-hardware/ https://blog.acolyer.org/2016/05/23/efficiently-compiling-ef...)
* CMU Peloton (http://db.cs.cmu.edu/papers/2017/p1-menon.pdf http://db.cs.cmu.edu/papers/2017/p1-menon.pdf)
* Vitesse (https://www.youtube.com/watch?v=PEmVuYjhQFo https://www.youtube.com/watch?v=PEmVuYjhQFo)
I think Greenplum was talking about doing this too, but that was about a year ago (http://engineering.pivotal.io/post/orca-profiling/ http://engineering.pivotal.io/post/orca-profiling/)
Most systems just compile the predicates (Impala, SparkSQL).
A lot of companies are talking about adding this now. The performance gains are significant.
- nimish 9y agoSql server does in memory compilation as well.
- pjmlp 9y agoJust as Oracle, using an intermediate representation originally designed for Ada compilers.
- assface 9y agoAFAIK, SQL Server's Hekaton engine only compiles the queries inside of stored procedures and UDFs. They do not compile every query that comes in from the client: https://docs.microsoft.com/en-us/sql/relational-databases/in-memory-oltp/survey-of-initial-areas-in-in-memory-oltp https://docs.microsoft.com/en-us/sql/relational-databases/in...
- mpala 9y agoThat's pretty misleading. All systems I am aware of essentially try to compile query-specific code and avoid re-compiling runtime code that doesn't vary between queries (e.g. if you look at the HyPeR paper, that's exactly what they describe). Compiling everything is questionable. There's not much point re-compiling runtime code or code outside of the hot path, it's expensive and doesn't bring any benefit. E.g. things that aren't beneficial to compile per query include: * Loops over a column of the same datatype, with no query specific branches (e.g. decoding a column of integers) * Other static code, e.g. some hash table operations * Outer loops that don't execute frequently * Rarely executed code, e.g. error handling. There are two general designs that let you compile only the necessary things. 1) the runtime calls into compiled code for hot loops vs 2) the compiled code drives the query and calls into the runtime. A lot of the systems you mentioned do the second, but Impala does the first, which seems to be the source of some misunderstanding. Also there were some cases where hot loops in earlier versions of Impala weren't compiled, but that's changed - generally we try to ensure that all hot loops are compiled. I think generally the optimal design wouldn't be complete query compilation, but rather something more like a traditional JIT that selectively compiles parts of the query. Source: work on Impala's query compilation
- jbapple 9y agoFWIW, Impala compiles more than just predicates. See, for instance, https://github.com/apache/impala/blob/e98d2f1c0af270930cd8a50d31b8924ad263937d/be/src/exec/hash-table.cc#L639 https://github.com/apache/impala/blob/e98d2f1c0af270930cd8a5...
- alecco 9y agoSQL query compilation in LLVM and no mention of HyperDB?? Edit: down votes... hah
- assface 9y agoWhich "HyperDB" are you talking about? I already listed Tableau's HyPer. Here are the other "hyper" systems that I know about and they don't do LLVM compilation: * HyperDex (http://hyperdex.org/ http://hyperdex.org/) * HyperGraphDB (http://www.hypergraphdb.org/ http://www.hypergraphdb.org/) * HyperSQL (http://hsqldb.org/ http://hsqldb.org/) * Hypertable (http://hypertable.org/ http://hypertable.org/)
- alecco 9y agoThe one from the groundbreaking 2011 VLDB paper on SQL query compilation with LLVM http://hyper-db.de http://hyper-db.de