3 ms·
the design you propose is to write 10sec chunks of data to an append-only log, which will contain chunks from multiple tables/series? if so, then reading larger
by jstrong 5y ago
the design you propose is to write 10sec chunks of data to an append-only log, which will contain chunks from multiple tables/series? if so, then reading larger ranges than 10sec requires scattered reads of a bunch of small chunks interspersed in the log file. that doesn't sound like it will be fast? the other big tradeoff is compression ratio, which writing data in 10sec chunks (as its final form) would be terrible for.
- IgorPartola 5y agoHow you actually write data to the files is mostly irrelevant. You could write each column of each table into a separate file or a file per table. I would probably further partition each file by date with some relatively large range (day or week or month). I don’t see why you’d ever want to write multiple tables into one file. Not sure why you’d think of this as chunks of 10 seconds, that’s just the size of the input buffer. The records aren’t related to each other beyond that or grouped in any way. As far as compaction, I suspect for the kinds of workloads you are using this kind of storage for you won’t get much in terms of deduplication. IoT devices sending sensor readings are unlikely to produce loads of stable readings and the time stamps will be ever increasing.