3 ms·
I’ve had some bad luck with hbase for similar but not tried spark - it’s better optimized for writers than readers iirc. Pinterest[1] did an interesting thing w
by idunno246 6y ago
I’ve had some bad luck with hbase for similar but not tried spark - it’s better optimized for writers than readers iirc. Pinterest[1] did an interesting thing where they worked directly on the hfiles. Everywhere I’ve worked I feel like I could pick any random data processing job that took more than a few minutes and get orders of magnitude improvement, there’s just so much inefficiency
[1]https://medium.com/pinterest-engineering/open-sourcing-terrapin-a-serving-system-for-batch-generated-data-7aa2f38c4472 https://medium.com/pinterest-engineering/open-sourcing-terra...
- anshumaniax 6y agoWhat do you mean inefficiency. We use the bulk upload feature and do billions of puts in an hour and our scans can go against 3 billion rows an hour. HBase scales linearly and we are already operating it on 5 times what we had designed it for
- idunno246 6y agoAh I meant that as separate comment, not hbase specifically, but that data pipelines need updating over time. Generally since what’s worth optimizing for changes as size increase, and there’s always that years old pipeline that takes hours to run that with a few changes could be minutes