3 ms·
The problem is that the initial writing is already so expensive, I guess we'd have to write multiple sensors into the same file instead of having one file per s
by Gasp0de 2y ago
The problem is that the initial writing is already so expensive, I guess we'd have to write multiple sensors into the same file instead of having one file per sensor per interval. I'll look into parquet access options, if we could write 10k sensors into one file but still read a single sensor from that file that could work.
- spothedog1 2y agoNew S3 Table Buckets [1] do automatic compaction [1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tab...
- necrobrit 2y agoTable buckets are currently quite hard to use for a lot of use cases as they _only_ support primitive types. No nested types. Hopefully this will come at some point. Product looks very cool otherwise.
- bloomingkales 2y agoSomething like Redis instead? [sensorid-timerange] = value. Your key is [sensorid-timerange] to get the values for that sensor and that time range. No more files. You might be able to avoid per usage pricing just by hosting this on a regular vps.
- Gasp0de 2y agoWe use Redis for buffering for a certain timeperiod, and then we write data for one sensor for that period to S3. However we fill up large Redis clusters pretty fast, so we can only buffer for a shortish period.
- hendiatris 2y agoYou may be able to get close with sufficiently small row groups, but you will have to do some tests. You can do this in a few hours of work, by taking some sensor data, sorting it by the identifier and then writing it to parquet with one row group per sensor. You can do this with the ParquetWriter class in PyArrow, or something else that allows you fine grained control of how the file is written. I just checked and saw that you can have around 7 million row groups per file, so you should be fine. Then spin up duckdb and do some performance tests. I’m not sure this will work, there is some overheard with reading parquet, which is why it is discouraged to have small files and row groups.
- 0cf8612b2e1e 2y agoWhy can you not rewrite the initial file into something partitioned by sensor+time? Would the one time job really be that much more additional cost vs the additional complexity of multiple sensors per file? Do you ever go back and reaggregate older data into bigger, sorted files? That is, maybe you originally partitioned by hour, but stale data is so infrequently accessed, you could roll up into partitions per week/month/whatever. Depending on the specifics, you might save some space from less file overhead and better compression statistics.
- Gasp0de 2y agoThe costly thing is the intial writing already. S3 is our cold storage, we don't often read from it. So compaction would only make reading cheaper, but create a writing cost in the process.