3 ms·
While I always love to see these kind of storage optimization I find the lack of metrics disappointing. Some metrics for cost saving vs query time increase wou
by TheGuyWhoCodes 4y ago
While I always love to see these kind of storage optimization I find the lack of metrics disappointing.
Some metrics for cost saving vs query time increase would be nice (taking into account query time range and partitioning time range i.e. number of files).
From a technical stand point when looking at S3 there is a minimum request first-byte time of 200ms and if you have multiple files to query on top of the list requests for the files in a bucket paths it can add up (you can save file/index metadata data in cache to help with this).
if the query is run once a day and isn't very latency sensitive I guess it doesn't matter but some solution that looks on the query data patterns and pre-loads data from s3 to hot/local cache storage or saves the file locally on the node for a period of time would surly help.
One scenario for this is ML models refresh of time series data where you might want to add new data with historical data and create a new model (even if it's incremental you'd still want to do regression testing with historical data). Of course there are code optimizations for these without the need of the DB to the the heavy lifting but that's just one thing that comes to mind.