3 ms·
Cloudera is doing just that with the recent announcement / open sourcing of Impala. Based on Amazon's description of their hosted product, the technology is ver
by monstrado 14y ago
Cloudera is doing just that with the recent announcement / open sourcing of Impala. Based on Amazon's description of their hosted product, the technology is very similar. Impala is still in beta, and columnar storage (trevni/avro) is right around the corner...with that, you can do petabyte scale queries for a very low cost.
https://github.com/cloudera/impala https://github.com/cloudera/impala
- flanger 14y agoPlatfora is doing some interesting work with interactive, in-memory BI for Hadoop. They essentially do away with the traditional DW/ETL model and create ephemeral in-memory 'lenses' for querying and visualization.
- capkutay 14y agoImpala is married to Hadoop. What if your data infrastructure isn't built on hbase and too complex/large to integrate it easily? Would impala still serve that purpose?
- TallGuyShort 14y agoIf your data is too difficult to integrate into HDFS (doesn't have to be HBase) using existing Hadoop tools, I suspect you're going to have to do some work to use that data on any platform.
- monstrado 14y agoImpala doesn't require HBase to operate, it can use raw HDFS. Simple example, if you had a few terabytes of TSV files, you could easily copy the raw data into HDFS and then create a simple schema around it. All queries on this data would be in parallel across all the nodes in the cluster, this is partly due to the distributed nature of HDFS.