3 ms·
This is pretty cool tech. I'd like to see this contrasted with the hot warm cold architecture which is the common approach to this problem. Usually people dea
by andrewvc 8y ago
This is pretty cool tech.
I'd like to see this contrasted with the hot warm cold architecture which is the common approach to this problem.
Usually people deal with this problem by allocating faster/more/less storage dense hardware for recent data and slower/less/more storage dense hardware for older data. You can read more about this here https://www.elastic.co/blog/hot-warm-architecture-in-elasticsearch-5-x https://www.elastic.co/blog/hot-warm-architecture-in-elastic...
Maybe something about it didn't work for them, but it's not clear to me from the article.
Disclaimer: I work for Elastic.
- karlney 8y agoWe tried the hot/cold architecture as well, and used it before in our data centre architecture. But currently we have concluded that it is not worth the added complexity. But that might change again if/when we learn more or get new requirements. We do change the number of replicas for recent data vs old data. We actually have a 4 tiered approach to how many replicas we use.
- Serow225 8y agoThanks for the article! We have a tiny hosted ELK cluster for our app logging, but we don't have much expertise in good cluster design. We recently implemented monthly rolling indices, and currently have a single 'archive' index for the old data (from the past two years) sharded into 25GB chunks. Would it make sense from a query performance perspective to add extra replicas to the 'archive' shards, to help spread the load between the nodes? Thanks for any thoughts!
- mikljohansson 8y agoA hot/cold tier architecture might work for setups that have a predictable and sharp cutoff from hot to cold. E.g. a logging setup, where the current days index gets all the indexing and queries. And there's a low concurrency occasional query for cold data Our workload is rather different from this. Indexing and document updates occur in an somewhat exponential decay pattern back in time, same with queries. So there's less sharp cut offs If we ran a hot-cold architecture we'd get a few issues Within each tier we'd get imbalanced workload over the nodes. Since within the tier the workload varies greatly with age of the indexes We use AWS i3 NVMe SSD instances. d2's with HDDs or using EBS have too long IO latency/throughput/iops even for our "cold" data workload. So a cold tier would scale based on storage needs, but in this tier we'd be wasting lots of compute capacity. And a hot tier would scale based on compute needs, and waste tons of storage capacity. By running both hot and cold workloads on the same set of nodes we get much more cost effective utilization. Since the hot workload uses most of the aggregate compute capacity, and the cold data uses most of the storage. But this then necessitates using Shardonnay to ensure we spread workload optimally across the clusters. And the more evenly we can spread it, the higher total utilization we can put on the clusters without having single nodes overload. A hot/cold architecture would much more costly for our workload. Since we'd have to unused storage on the hot tier, and unused compute capacity on the cold tier. A single tier just makes much more sense for our particular use case
- outworlder 8y agoHot/Warm architecture is nice, specially for logging setups. But it requires quite a bit of supporting logic (usually scripts) which is not available on ES itself.