4 ms·
It's interesting to see the continued trends to separate compute from storage. I'm curious about the tradeoffs. The article mentions increased latency of 10's
by d_watt 4y ago
It's interesting to see the continued trends to separate compute from storage.
I'm curious about the tradeoffs. The article mentions increased latency of 10's of milliseconds. If you have a complex access (many joins or index hits), is that exacerbated from many round trips to load different files? From a cost perspective, I image that you now have to think about IOPS. If you have a very high usage database, would it make sense to have a more traditional file system where you aren't paying per operation?
Edit: On reread, I missed that this is only available for the hypertable concept, and is more of a "cold storage" for older metrics, presumably accessed infrequently. I am curious about a general "PG backed by S3"'s ramifications.
- mfreed 4y agoThis is a great observation. As you point out, this was designed for the workload patterns we typically see with time-series, events, and analytical data, where the query (& insert) patterns differ across time. So I agree that it's good for cold storage, but it's a bit nuanced. For example, you rarely see small random queries to old historical data, but you do often see larger scans over historical data. And in those cases, the throughput you get from S3 is actually quite high (especially that we've engineered with with proper columnar compression and row group/columnar exclusion). Which is very different from a latency bounded workload where you have a lot of small random reads, which is much more common in CRUD-like workloads. Also, with Timescale, you have the ability to build continuous aggregates (incrementally materialized views). So you can have the raw data (or even lower levels of rollups) that get tiered into S3, while the more frequently accessed rollups can remain in hot storage. (Timescale co-founder)
- leoalves 4y agoAnother option as a "PG backed by S3" is neon.tech