3 ms·
Presto does read Hive data directly out of HDFS, bypassing Hive/MapReduce. However, neither HDFS nor the file formats like RCFile are efficient. The entire desi
by electrum 13y ago
Presto does read Hive data directly out of HDFS, bypassing Hive/MapReduce. However, neither HDFS nor the file formats like RCFile are efficient. The entire design of HDFS makes it difficult to build a real column store. Having a native store for Presto allows us to have tight control over the exact data format and placement, enabling 10-100x performance increases over the 10x that Presto already has over Hive.
Additionally, in the long term, we want to enable Presto to be a completely standalone system that is not dependent on HDFS or the Hive metastore, while enabling next-generation features such as full transaction support, writable snapshots, tiered storage, etc.
- flyovercountry 13y agoHDFS now supports placement groups with one use case being columnar storage.