3 ms·
No joins. I wonder why you wouldn't use PrestoDB to connect to Elastic Search. It provides you with an SQL engine and you just need to write a connector that k
by jermo 11y ago
No joins.
I wonder why you wouldn't use PrestoDB to connect to Elastic Search. It provides you with an SQL engine and you just need to write a connector that knows how to get data.
Similar thing has been done in Crate.io.
- rmsaksida 11y agoIt does have joins. https://github.com/NLPchina/elasticsearch-sql/blob/5cd6ab63919bd4f366fc721638b31a54e3555b5c/src/test/java/org/nlpcn/es4sql/JoinTests.java https://github.com/NLPchina/elasticsearch-sql/blob/5cd6ab639...
- deleted 11y ago[deleted]
- lobster_johnson 11y agoInteresting — looks like the join isn't pipelined. The entire right-hand side evaluates synchronously. So it has to wait for the entire right-hand result set before it can evaluate the join operator, instead of streaming it concurrently. I'm surprised anyone would do it this way in Java, which has good support for concurrency. Edit: Actually the file you linked to was a test file. Hash join code is here [1], and it uses ES' scrolling feature to incrementally join, though it's not pipelined. Not sure scrolling is entirely appropriate for this; it will potentially hold an unpredictable amount of memory on the server end. [1] https://github.com/NLPchina/elasticsearch-sql/blob/5cd6ab63919bd4f366fc721638b31a54e3555b5c/src/main/java/org/elasticsearch/plugin/nlpcn/HashJoinElasticExecutor.java https://github.com/NLPchina/elasticsearch-sql/blob/5cd6ab639...
- gdulli 11y agoJoins are absent from the documented list of features and examples on the initial README but there's "limited support" for joins as described here: https://github.com/NLPchina/elasticsearch-sql/wiki/Join https://github.com/NLPchina/elasticsearch-sql/wiki/Join
- joefkelley 11y agoDon't have much experience with Presto, but I have used Hive to query Elastic Search. It works very very well for full-table-scan analytics-type queries. I expect PrestoDB would be similar. But when it comes to queries about smaller and smaller pieces of the full dataset, it becomes less and less likely that these types of connectors perform well. Predicate pushdown is rarely well-implemented in these types of "run SQL against any big data" systems (Hive, Presto, Impala, SparkSQL, etc). A simple "select * where id = 1234" will often do a full scan and filter within the query engine, rather than push the point lookup into ES.
- rxin 11y agoActually Spark SQL's data source API has a very expressive predicate pushdown interface and most data sources implement them. id = 1234 should not do a full scan.
- threeseed 11y agoAs mentioned Spark SQL has good predicate push down support. The ElasticSearch, Cassandra, HBase, MongoDB adapters all support it.