3 ms·
That is correct, we are using the functionality similar to HBase. We pull in the HBase BlockCache library with some tweaks to make it work for our scenario. Not
by varunsharma 11y ago
That is correct, we are using the functionality similar to HBase. We pull in the HBase BlockCache library with some tweaks to make it work for our scenario. Note that data is never overwritten and HFiles are immutable. So the cache automatically, gets evicted/populated as HFiles are opened and closed. That said, there is a possibility to use more performant formats like rocksdb etc. in the future (the format is pluggable). Or even still use HFiles and have them loaded into some kind of specialized in memory data structure etc.
- ameyamk 11y agoCool. So in a way its Immutable read only HBase (with guaranteed data locality no memstore overhead and compactions overhead). Cool. Nice solution. I wonder this can be patched back to HBase - as a "read only mode" ?
- varunsharma 11y agoSuspect it would be difficult. There are differences like new data is completely independent of previous data and the source of truth for region distribution is HDFS for the block locations. There are multiple replicas per shard while HBase has one region for each shard etc. Across bulk loads, the number of shards (regions) can be changed for the same fileset or table - not possible with HBase. The data sharding is not range based like in HBase but mod based, as output by Hadoop HashPartitioner. There are so many differences that its hard to accommodate it into the HBase code.