3 ms·
There are some unique challenges to storing time series data that are different than those of relational databases. Namely, read/write asymmetry, data safety, d
by camel_gopher 8y ago
There are some unique challenges to storing time series data that are different than those of relational databases. Namely, read/write asymmetry, data safety, data aggregation, and analysis of large data sets.
I wrote in depth about these problems and how different TSDBs solve them here.
https://www.irondb.io/2018/08/tsdbs-at-scale-part-one/ https://www.irondb.io/2018/08/tsdbs-at-scale-part-one/
- manigandham 8y agoAll modern columnstores can handle vast ingest rates and query speeds. It's all down to sharding, zone maps and sparse indexing, fast algorithms that operate on compressed data, and storage throughput. These are well-solved problems at this point. Your blog post doesn't mention a single columnstore database though. KDB+, Clickhouse, MemSQL, or any of the GPU-powered variations will happily beat any TSDB out there.
- camel_gopher 8y agoSure they can handle them, just not in an economically viable fashion.
- manigandham 8y agoCompared to what? Economically viable is very vague and relative. Columnar storage can easily reach 90% compression levels, is faster to read, and vectorized processing beats per-row/record iteration, so there's a reason it's the best for OLAP currently. Why not benchmark IronDB against Clickhouse and post the results?
- chaotic-good 8y ago90% compression on time series is not viable unless you have some very specific dataset.
- manigandham 8y agoBoth of you commenters have your own TSDBs which seems to be coloring all of your posts. I'm going to leave this conversation as unproductive unless you care to benchmark your products against modern column-stores, although I think it's telling that there are never such benchmarks available.
- chaotic-good 8y agoThey can't handle high cardinality. Imagine having millions of columns in the column-oriented database (70% of those columns are updated every second). Imagine that you have to add new columns all the time. The main misconception about TSDB's is that it's just a data with timestamp. TSDB's has multi-dimentional data model, time is only one of the dimensions.
- manigandham 8y agoYou don't need to add new columns. CREATE TABLE metrics (metric_name text, ts timestamp, properties json, key(metric_name, ts)) OLAP queries with SQL are very good at handling whatever dimensions you want.
- chaotic-good 8y ago'metric_name text' is actually a tag-value list. Many TSDB's allows you to match data by tag. Each tag should be represented by a column in your example. Single table design will be prone to high read/write amplification due to data alignment. Usually, you need to read many series at once so your query will turn into full table scan. Or it will read a lot of unneeded data which happened to be located near the data you need. Writes will be slow since your key starts with metrics name. Imagine that you have 1M series and each series gets new data point every second. In your scema it will result in 1M random writes. Cardinality of the table will go through the roof, BTW. Every data point will add the key. Good luck dealing with this.
- manigandham 8y agoYou are talking about regular relational databases. I'm talking about distributed column-oriented databases. Big difference. You can store tags and other data in JSON/ARRAY columns. The primary key is used for automatically sharding and sorting. Groups of rows are sorted, split into columns, compressed, and stored as partitions with metadata. This means you can 'scan' the entire table in milliseconds using metadata and then only open the partitions, and the columns inside, that you actually need for your query. There are no random writes either, it's all constant sequential I/O with optional background optimization. And because of compression, storing the same key millions of times has no real overhead. As stated several times before, we deal with this everyday on trillion row tables inserting 100s of billions of rows daily. Queries run in seconds. We do just fine.