4 ms·
I'm glad more people are tackling this problem. There still isn't a good solution to real-time aggregation data at large scale. At a previous company, we dealt
by temuze 6y ago
I'm glad more people are tackling this problem. There still isn't a good solution to real-time aggregation data at large scale.
At a previous company, we dealt with huge data streams (~1TB data / minute) and our customers expected real-time aggregations.
Making an in-house solution for this was incredibly difficult because each customer's data differed wildly. For example:
- Customer A's shards might have so much cardinality where memory becomes an issue.
- Customer B's shards might have so much throughput where CPU becomes a constraint. Sometimes a single aggregation may have so much throughput where you need to artificially increase the cardinality and aggregate the aggregations!
This makes the optimal sharding strategy very complex. Ideally, you want to bin-pack memory-constrained aggregations with CPU-constrained aggregations. In my opinion, the ideal approach involves detecting the cardinality of each shard and bin-packing them.
- jstrong 6y agoI've always found that when you are solving a concrete problem, like you were, it's vastly easier than the case of a general-purpose database because you can make all the tradeoffs that benefit your exact use case. but it sounds like that's not what you experienced. was it just how heterogeneous the clients' needs were? I guess what I'm saying is, if you are capable of handling 1TB/minute, seems like you're plenty able to and would want to be designing the system yourself - but interested what I'm missing about this.