3 ms·
> you prevent the database from growing without bound? The whole point of shipping cold storage off to S3 was to solve that kind of scale problem. You can ke
by gopalv 5y ago
> you prevent the database from growing without bound?
The whole point of shipping cold storage off to S3 was to solve that kind of scale problem.
You can keep your ingest nodes scaled up to the incoming data (per-day, approx), replicate 3-way for HA and use SSDs for commit throughput.
The query nodes scaled up to the working set sizes, but auto-scale up/down based on the workload scan sizes (no need to keep them running 24x7, cheaper to throw away the cache after a workday - no need for replicas, just jitter them so that the entire cache doesn't go poof at the same time + hit s3 throttling on the next query).
And the S3 bucket is literally infinite storage (more like caps out when the metadata about the s3 backed items hits 8 Tb).
Run the equivalent of an fsck every quarter to check the checksums, then re-encrypt them with a new key or to recompress the blocks by swapping them (go from lz4 to zstd as they age out).
There is a mechanism to expire data after 7 years of storage (I guess it won't be queried anymore, so there'd be nothing "live" to expire?), but that might be longer than the current architecture lives on without a refactor.