3 ms·
Suspect it would be difficult. There are differences like new data is completely independent of previous data and the source of truth for region distribution i
by varunsharma 11y ago
Suspect it would be difficult. There are differences like new data is completely independent of previous data and the source of truth for region distribution is HDFS for the block locations. There are multiple replicas per shard while HBase has one region for each shard etc. Across bulk loads, the number of shards (regions) can be changed for the same fileset or table - not possible with HBase. The data sharding is not range based like in HBase but mod based, as output by Hadoop HashPartitioner. There are so many differences that its hard to accommodate it into the HBase code.