4 ms·
(Disclaimer: post author and Timescale employee) I'm sorry you feel like we were trying to be dishonest in the post. On the contrary, we put a lot of effort (a
by ryanbooz 6y ago
(Disclaimer: post author and Timescale employee)
I'm sorry you feel like we were trying to be dishonest in the post. On the contrary, we put a lot of effort (and 7,000+ words) into trying to explain everything that we did - just as we've done with other benchmarks which others have linked to.
The TimescaleDB test did not use continuous aggregates for these test, only raw time-series data stored in hypertables.
For each database, we (and other contributors) do our best to use features in all cases that take advantage of the DB. For TimescaleDB, a function like LAST() happens to be really powerful for most workloads and is really, really fast. That's not cheating, it's using the software properly! :-)
The SQL that we generate for each database can be examined here (https://github.com/timescale/tsbs/tree/master/cmd/tsbs_generate_queries/databases https://github.com/timescale/tsbs/tree/master/cmd/tsbs_gener...) and as an open source project, anyone is free to contribute!
- BenoitP 6y agoSorry for the harsh comment. I've been reading about your columnar compression pipeline [1], and it sort of makes sense if the comparison is against a regular row-oriented DB. AWS Timestream must really be doing something wrong here, or serving an entirely different use case. 5-175x faster queries and 150x-220x cheaper I do get it. But 6000x higher inserts does not make sense to me. It is insane, and literally unbelievable to me. Storage savings are at 96% for "IT metrics (DevOps dataset from TSBS)", so it should be closer to 25x higher insert rate. Where is the missing 240x? Is this from some distributed replication overhead? Is this from local vs remote insertion? Is this from bulk inserts vs per row? Anyway I wanted to thank you for your kind efforts in writing the blog post and providing answers here; and for the patience that you show to the audience here, me included. [1] https://blog.timescale.com/blog/building-columnar-compression-in-a-row-oriented-database/ https://blog.timescale.com/blog/building-columnar-compressio...
- ryanbooz 6y agoWe're happy for people to poke at this and helping us to improve. It's obviously hard to work at something for weeks, see the numbers (even knowing you really tried for days to move the needle) and then still publish numbers that seem impossible. And again, if you look at TSBS, this isn't the first time we've run benchmarks on other databases, so we were just as shocked and put extra effort into it. In the end, if you read the article (and not just the headlines - not saying you are, but it's easy to see 6000x and latch on to it), the comparison is absolutely focused on this one, pretty straight forward use case (although we normally run 5 different scenarios): From one client, given a specific kind of workload (100 hosts, 10 CPU metrics every 10 seconds for 30 days = ~1 billion metrics) - how fast could we save the data. Most other time series databases at least perform marginally well with the same setup... load data with one client. But Timestream just doesn't seem setup to work that way. Some of the responses today imply that we need really large clients with thousands of threads to get those speeds. And that might work if we kept going and spent more time and significantly more money. We just haven't ever had to do that before. If your use case better aligns with what Timestream offers, then it might be a great product for you. Given some of the many other concerns we discovered along the way, it doesn't yet seem like the time to jump in. All the best!
- varjoranta 6y agoAmazon Timestream uses quorum writes as an example. https://aws.amazon.com/blogs/aws/store-and-access-time-series-data-at-any-scale-with-amazon-timestream-now-generally-available/ https://aws.amazon.com/blogs/aws/store-and-access-time-serie... "This is made possible by the way Timestream is managing data: recent data is kept in memory and historical data is moved to cost-optimized storage based on a retention policy you define. All data is always automatically replicated across multiple availability zones (AZ) in the same AWS region. New data is written to the memory store, where data is replicated across three AZs before returning success of the operation. Data replication is quorum based such that the loss of nodes, or an entire AZ, does not disrupt durability or availability. In addition, data in the memory store is continuously backed up to Amazon Simple Storage Service (S3) as an extra precaution." Feels like apples to oranges comparison, as the consistency models are really different. Then again, I woul definitely optimize for performance on most timeseries use cases. Different products with different features baked in.