5 ms·
Will be very interesting if someone does a benchmark comparison of Presto with Cloudera Impala, Amazon RedShift and Apache Drill. Also, very curious to know (f
by dude_abides 13y ago
Will be very interesting if someone does a benchmark comparison of Presto with Cloudera Impala, Amazon RedShift and Apache Drill.
Also, very curious to know (from any Googlers browsing HN) if Dremel is still the state-of-the-art within Google, or if there is already a newer replacement.
- snorkel 13y agoI can least say that I've used both a large Hadoop+Hive farm and a moderately sized RedShift cluster on the same large data set, and holy smoke, RedShift is orders of magnitude faster. Results vary by how big the nodes are you allocate to Redshift, and it's not cheap at all, but very impressive.
- dman 13y agoI am working on a related product. Is there some way I could contact you to pick your brain about some things?
- spikels 13y agoThis is exactly what you should expect in almost all cases because you are kinda comparing apples and oranges. Hadoop/Hive was not focused on speed but proving it is even possible to do queries reliably on such large datasets. Once this was achieved it immediately became clear the long waits for Hadoop/Hive batch jobs to finish made it impractical for many uses. Presto, Impala, Drill, RedShift, etc were all designed Primarily to address this problem and be much, much faster than Hadoop/Hive so that the data could be queried interactively. All these new projects/products are in a very active competition to find the best way or ways to do this. You should compare RedShift to these other projects rather than Hadoop if speed is an issue for you. Each has it's pluses and minuses depending on the situation.
- aquadrop 13y agoWait, so RedShift (Drill, Impala, Presto) do the same job as Hadoop\Hive, only faster? You only mentioned they addressed the speed, but what's downside? At what cost they achieved their velocity?
- necubi 13y agoBasically, Hive is incredibly inefficient. Hive works by taking apart the HiveQL query and turning it into a series of Hadoop MR steps. A complex hive query may have many such steps, and for each one data must be read from HDFS, processed, and written back to HDFS. For most queries, the vast majority of your time will be spent just doing IO. Hadoop is also very, very slow to start up tasks (> 1 minute), so when you have a lot of them that can come to dominate your total run time. Impala, Drill, etc. avoid all those unnecessary reads and writes by implementing the querying logic directly, rather than by compiling to Map Reduce. Shark [0] is an interesting counterpoint. It takes essentially the same approach as Hive but on Spark instead of Hadoop, and achieves similar or better performance than the more "direct" implementations. [0] http://shark.cs.berkeley.edu http://shark.cs.berkeley.edu
- spikels 13y agoThe tradeoffs are both complex and evolving and are different for each project/product. In general Hadoop is more scalable and more flexible but much slower. For example as of today RedShift can hold a maximum of 256 terrabytes of compressed data while Facebook's Hadoop cluster was over 200 Petabytes in late 2012. RedShift only supports limited query and data types and a single index while Hadoop can theoretically handle arbitrary data processing. But if these constraints are acceptable then RedShift will likely be orders of magnitude faster in most cases. Other projects/products will have different tradeoffs but they are almost always faster as this was almost always the primary goal.
- cjg_ 13y agoWhere do you get the limit of 256TB compressed from? Amazon Redshift enables you to start with as little as a single 2TB XL node and scale up all the way to a hundred 16TB 8XL nodes for 1.6PB of compressed user data. from http://aws.amazon.com/redshift/features-and-benefits/ http://aws.amazon.com/redshift/features-and-benefits/
- spikels 13y agoMy mistake, thank for the correction. The 16 node limit I was familiar with is not a hard limit. You can request more nodes. http://aws.amazon.com/redshift/faqs/#0080 http://aws.amazon.com/redshift/faqs/#0080 It would be interesting to see performance comparisons on these huge datasets. I would expect to see new and interesting problems at that scale.
- necubi 13y agoThe state of the art at Google appears to be F1 [0], although that takes a pretty different approach than Presto/Dremel and may be complementary. This area (interactive SQL on big data) has become very active lately. In addition to those you list, there's Shark [1] and BlinkDB [2] from Berkeley's AMP Lab. [0] http://research.google.com/pubs/pub41344.html http://research.google.com/pubs/pub41344.html [1] http://shark.cs.berkeley.edu http://shark.cs.berkeley.edu [2] http://blinkdb.org/ http://blinkdb.org/
- sprizzle 13y agoI believe F1 is the SQL DB that powers ads (AdWords, AdSense, etc.) and is used in production for consumer-facing apps. Dremel is more of a data/log analysis tool, but doesn't typically interact directly with consumer-facing apps. Dremel (externally: Google BigQuery) is still widely used at Google across all product areas; F1 is used in a few products but not many.
- snewman 13y agoAs sprizzle says, F1 and Dremel are not competitors; Dremel is a read-only system supporting large analysis queries, while F1 is a read-write system designed for large numbers of smaller queries. Last year Google published a paper on PowerDrill (http://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf http://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf), which sounds somewhat like Dremel but is designed to use large amounts of RAM, which in turn enables some very powerful optimizations.
- matsur 13y agoWould also be curious to see how the execution model and language relates to/differs from Microsoft Dryad[0]/SCOPE[1]. [0] http://research.microsoft.com/en-us/projects/dryad/ http://research.microsoft.com/en-us/projects/dryad/ [1] http://research.microsoft.com/en-us/um/people/jrzhou/pub/Scope.pdf http://research.microsoft.com/en-us/um/people/jrzhou/pub/Sco...
- curiousfiddler 13y agoYou might want to add Apache Tez to the list as well. Series of posts on Tez: http://hortonworks.com/hadoop/tez/ http://hortonworks.com/hadoop/tez/
- tonfa 13y agoFor interactive workflows, there is powerdrill: http://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf http://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf
- monstrado 13y agoAlthough a little out of date, there is a website dedicated to this: https://amplab.cs.berkeley.edu/benchmark/ https://amplab.cs.berkeley.edu/benchmark/
- dude_abides 13y agoThis is awesome, thanks for sharing! Redshift looks to be order-of-magnitude faster than Impala or Shark in all the test. Does this mean that once RedShift supports user-defined functions, there is no competing solution that is any match? (Unless you want to avoid using the cloud)
- flyovercountry 13y agoThe file format for hadoop tools is a sequence file. For most queries this is the second slowest format, after text. Rcfile or parquet would be a more interesting benchmark.
- monstrado 13y agoThe benchmark above is testing Impala with SequenceFiles compressed with GZIP, against RedShift, which is not a fair comparison. In the "What's next?" section, they say they want to re-do the Impala tests using Parquet, which is a columnar format based on the Dremel whitepaper (http://parquet.io/ http://parquet.io/).
- dude_abides 13y agoAh that makes sense! Looking at their results and how RedShift was so much faster in every scenario, it looked like something was amiss. Is Parquet Cloudera-only like Impala or is it available with vanilla Hadoop?
- monstrado 13y agoImpala isn't technically Cloudera only, it's open source (https://github.com/cloudera/impala https://github.com/cloudera/impala), and other people have gotten it to run on their Hadoop distribution, but since it's developed by Cloudera, it was developed to run on the CDH platform (Hadoop). Parquet was a joint effort between Cloudera and Twitter, and now it's being developed by many other companies. You can use it with Hive, Pig, MapReduce, Cascading, Crunch and I think Apache Drill's first milestone has adopted it as a columnar format as well. Parquet also allows you to use your Avro or Thrift schema (soon Protobuffs) to write Parquet data, too. It's a separate project in the ecosystem and has its own roadmap (https://github.com/Parquet/parquet-mr https://github.com/Parquet/parquet-mr).
- spikels 13y agoThere is also Shard-Query which is build on top of a cluster on MySQL servers. According to its developer, who works at Percona, running Shard-Query with MySQL using a column store (Infobrite community edition) give performance similar to RedShift. http://shardquery.com/ http://shardquery.com/
- justinerickson 13y agoFor those of you interested in more information on Impala and it's performance characteristics see: * rideimpala.com * https://speakerdeck.com/grahn/practical-performance-analysis-and-tuning-for-cloudera-impala https://speakerdeck.com/grahn/practical-performance-analysis...