8 ms·
Echoing the sentiment expressed by others here, for a scalable time-series database that continues to invest in its community and plays well with others, please
by avthar 6y ago
Echoing the sentiment expressed by others here, for a scalable time-series database that continues to invest in its community and plays well with others, please check out TimescaleDB.
We (I work at TimescaleDB) recently announced that multi-node TimescaleDB will be available for free, specifically as a way to keep investing in our community: https://blog.timescale.com/blog/multi-node-petabyte-scale-time-series-database-postgresql-free-tsdb/ https://blog.timescale.com/blog/multi-node-petabyte-scale-ti...
Today TimescaleDB outperforms InfluxDB across almost all dimensions (credit goes to our database team!), especially for high-cardinality workloads:
https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-time-series-data-timescale-influx-sql-nosql-36489299877/ https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-...
TimescaleDB also works with Grafana, Prometheus, Telegraf, Kafka, Apache Spark, Tableau, Django, Rails, anything that speaks SQL...
- jimaek 6y agoWhile influx is pretty bad overall it's super simple to deploy and configure unlike timescale. It's the main reason we decided to use influx in our small team with simple enough timeseries needs
- avthar 6y agoCurious what you found difficult to deploy/ configure? Is this in a self-managed context?
- samstave 6y agoI found this to be a funny, subtle insult. :-) "Why is it difficult, because you're self managed?"
- themgt 6y agoI was testing out both TimescaleDB/InfluxDB recently, including maybe using with Prometheus/Grafana. I was leaning towards Timescale, but InfluxDB was indeed a lot easier to quickly boot a "batteries included" setup and start working with live data. I eventually spent a while reading about Timescale 1 vs 2, and testing the pg_prometheus[1] adapter and started thinking through integrating its schema to our other needs then realizing it's "sunsetted" and then reading about the new timescale-prometheus[2] adapter and reading through its ongoing design doc[3] with updated schema that I'm less a fan of. I finally wound up mostly-settling on Timescale although I've put the Prometheus extension question on hold, just pulling in metrics data and outputting with ChartJS and some basic queries got me a lot closer to done for now. Our use case may be a little odd regardless, but I think a timescale-prometheus extension with a some ability to customize how the data is persisted would be quite useful. [1] https://github.com/timescale/pg_prometheus https://github.com/timescale/pg_prometheus [2] https://github.com/timescale/timescale-prometheus https://github.com/timescale/timescale-prometheus [3] https://docs.google.com/document/d/1e3mAN3eHUpQ2JHDvnmkmn_9rFyqyYisIgdtgd3D1MHA/edit#heading=h.hv9q074m6qpz https://docs.google.com/document/d/1e3mAN3eHUpQ2JHDvnmkmn_9r...
- cevian 6y ago(TimescaleDB engineer) Really curious about finding out more about your reservations about the updated schema. All criticisms welcome.
- themgt 6y agoThanks, as I say our use case may be too odd to be worth supporting, but effectively we're trying to add a basic metrics (prom/ad-hoc) feature to an existing product (using Postgres) with an existing sort of opinionated "ORM"/toolkit for inserting/querying data. Because of that and the small scale required, the choice of table-per-metric would be a tough fit and I think a single table with JSONB and maybe some partial indexes is going to work a lot better for us. It would just be nice if we could somehow code in our schema mapping and use the supported extension, but I get it may be too baked-into the implementation. Anyway, overall we're quite happy with TimescaleDB!
- souldeux 6y agoWhat's wrong with influx? I use it and like it, albeit for hobby-level projects.
- mhall119 6y agoAs with any technology, the right tool for the job is highly dependent on what the job is.
- chucky_z 6y agoInflux is so good for hobby-level projects. So good!! For serious applications, it doesn't cut it. Trying to do more than a few TB a day is a waste of time outside of enterprise, which ain't cheap. I plopped VictoriaMetrics in place of Influx for my cases and haven't even had a single hiccup.
- valyala 6y agoInfluxDB may have slightly big RAM requirements when working with high number of time series [1]. [1] https://medium.com/@valyala/insert-benchmarks-with-inch-influxdb-vs-victoriametrics-e31a41ae2893 https://medium.com/@valyala/insert-benchmarks-with-inch-infl...
- heliodor 6y agoFor small and medium needs, InfluxDB is a delight to use. On top of that, the data exploration tool within Chronograf is one-of-a-kind when it comes to exploring your metrics. If you are looking for a vendor to host, manage, and provide attentive engineering support, check out my company: https://HostedMetrics.com https://HostedMetrics.com
- valyala 6y agoIf you want the best of both worlds, then try VictoriaMetrics. It is simple to deploy and operate and it is more resource-efficient comparing to InfluxDB. More on this, it supports data ingestion over Influx line protocol [1]. [1] https://victoriametrics.github.io/#how-to-send-data-from-influxdb-compatible-agents-such-as-telegraf https://victoriametrics.github.io/#how-to-send-data-from-inf...
- aeyes 6y agoI don't consider TimescaleDB to be a serious contender as long as I need a 2000 line script to install functions and views to have something essential for time-series data like dimensions: https://github.com/timescale/timescale-prometheus/blob/master/pkg/pgmodel/migrations/sql/1_base_schema.up.sql https://github.com/timescale/timescale-prometheus/blob/maste... https://github.com/timescale/timescale-prometheus/blob/master/extension/sql/timescale-prometheus.sql https://github.com/timescale/timescale-prometheus/blob/maste...
- gen220 6y ago(Not affiliated with TimescaleDB, just trying to understand your critique) TimescaleDB is a Postgres extension, which means they have to play PG's rules, which rather implies a complicated mesh of functions, views, and "internal-use-only" tables. But once they're there, you can pretty much pretend they don't exist (until they break, of course, but this is true for everything in your software stack). Is your complaint targeting the size of these scripts, or their contents? Because everything in there appears pretty reasonable to me.
- klohto 6y agoWhich kinda confirms the parent’s issues... InfluxDB is really easy to deploy and forget. With TimescaleDB you should be ready to know ins-and-outs of PG to secure and maintain correctly. Sure, for scaling and high loads TDB might be good but InfluxDB is easier and suitable for most loads and maintainability.
- gen220 6y ago> With TimescaleDB you should be ready to know ins-and-outs of PG to secure and maintain correctly. Gotcha, this makes sense. To this, I'd pose the question: is the same not true for Influx? (i.e. With IFDB, you should be ready to know its ins and outs to secure and maintain it correctly). I guess I think about choosing a database like buying a house. I want it to be as good in X years as it is today, maybe better. From this perspective, PG has been around for a long time, and will continue to be around for a long time. When it comes time to make tweaks to your 5 year old database system, it will be easier to find engineers with PG experience than Influx experience. Not to mention all of the security flaws that have been uncovered and solved by the PG community, that will need to be retrodden by InfluxDB etc. Anyways, it's just an opinion, and it's good to have a diversity of them. FWIW, I think your perspective is totally valid. It's always interesting to find points where reasonable people differ.
- gshulegaard 6y agoJust wanted to say I am super impressed with the work TimescaleDB has been doing. Previously at NGINX I was part of a team that built out a sharded timeseries database using Postgres 9.4. When I left it was ingesting ~2 TB worth of monitoring data a day (so not super large, but not trivial either). Currently I have built out a data warehouse using Postgres 11 and Citus. Only reason I didn't use TimescaleDB was lack of multi-node support in October of last year. I sort of view TimescaleDB as the next evolution of this style of Postgres scaling. I think in a year or so I will be very seriously looking at migrating to TimescaleDB, but for now Citus is adequate (with some rough edges) for our needs.
- roskilli 6y agoIf you're looking at scaling monitoring timeseries data you may also wanter to consider more Availability leaning architecture (in the CAP theory sense) with respect to replication (i.e. quorum write/read replication, strictly not leader/follower - active/passive architecture) then you might also want to check out the Apache 2 project M3 and M3DB at m3db.io. I am biased obviously as a contributor. Having said that I think it's always worth understanding active/passive type replication and the implications and see how other solutions handle this scaling and reliability problem to better understand the underlying challenges that will be faced with instance upgrades, failover and failures in a cluster.
- gshulegaard 6y agoNeat! Hadn't heard of M3DB before, but cursory poke around the docs seems like it's a pretty solid solution/approach. My current use case isn't monitoring, or even time series anymore, but will keep M3DB in mind next time I have to seriously push a time series/monitoring solution.
- valyala 6y agoDid you try ClickHouse? [1] We were successfully ingesting hundreds of billions of ad serving events per day to it. It is much faster at query speed than any Postgres-based database (for instance, it may scan tens of billions of rows per second on a single node). And it scales to many nodes. While it is possible to store monitoring data to ClickHouse, it may be non-trivial to set up. So we decided creating VictoriaMetrics [2]. It is built on design ideas from ClickHouse, so it features high performance additionally to ease of setup and operation. This is proved by publicly available case studies [3]. [1] https://clickhouse.tech/ https://clickhouse.tech/ [2] https://github.com/VictoriaMetrics/VictoriaMetrics/ https://github.com/VictoriaMetrics/VictoriaMetrics/ [3] https://victoriametrics.github.io/CaseStudies.html https://victoriametrics.github.io/CaseStudies.html
- pritambaral 6y agoObligatory heads-up: parts of TimescaleDB (including the multi-node feature) come with no right-to-repair. See https://news.ycombinator.com/item?id=23274509 https://news.ycombinator.com/item?id=23274509 for more details.
- akulkarni 6y agoWe are currently working on revising this to make the license more open. Stay tuned :-). (If you want to provide any early feedback, please email me at ajay (at) timescale.com)
- etxm 6y ago> high-cardinality workloads Does this mean if you were using it with Prometheus you could get around issues with high cardinality labels?
- say_it_as_it_is 6y agoThis thread is about infra monitoring, where any of the potential performance differences probably don't matter at all. If it does, show us. I don't work for any time series database provider.
- darkwater 6y agoWhy performances shouldn't matter in this case? The more instances, the more timeseries and the more datapoints you have, the more you care for speed when you need to query it to visualize for example WoW changes in, say, memory usage.
- say_it_as_it_is 6y agoConsider the case at hand, involving streaming charts and alerts. There will be zero perceptible difference in the streaming charts regardless of what database is used. Alerts won't trigger for whatever millisecond difference there may be, and I don't think that this matters to any developer or manager awaiting the 3am call.
- brown9-2 6y agoTime series database performance is not only about the reads.
- didip 6y agoHow does multi node works with Postgres? Does TimescaleDB create its own Raft layer on top and just treat Postgres as dumb storage?
- k-rus 6y agoA TimescaleDB engineer here. Current implementation of database distribution in TimescaleDB is centralised where all traffics go through an access node, which distributes the load into data nodes. The implementation uses 2PC. Abilities of PostgreSQL to generate distributed query plans are utilised together with TimescaleDB optimisations. So PostgreSQL is used not just a dumb storage :)
- richardARPANET 6y agoWe're using TimescaleDB + Grafana for visualising sensor data for a Health Tech product. No complaints so far.
- znpy 6y agoAfter a few years in the industry of systems engineering and administration I think that "no complaints so far" is one of the best compliments a software can receive.
- ngrilly 6y agoTimescaleDB and my team is using it, but one significant drawback compared to solutions like Prometheus are the limitations of continuous aggregations (basically no joins, no order by, no window functions). That’s a problem when you want to consolidate old data.
- cevian 6y ago(TimescaleDB engineer) we hear you and are working on making continuous aggregations easier to use. For now, the recommended approach is to perform continuous_aggregates on single tables and perform joins, order by, and window when querying the materialized aggregate rather than when materializing. This often has the added benefit of often making the materialization more general so that a wider range of queries can use it.
- ngrilly 6y agoThanks for the advice! Makes sense. We are doing something similar (aggregate on single table and join later). But still looking a solution to compute aggregated increments when there are counter resets.
- ngrilly 6y agoMeant "TimescaleDB is great and my team is using it"...
- mekster 6y agoCan you refute the claim that TimescaleDB uses such a huge disk space that is 50 times larger than other time series database? I see no reason to use it over VictoriaMetrics. https://medium.com/@valyala/high-cardinality-tsdb-benchmarks-victoriametrics-vs-timescaledb-vs-influxdb-13e6ee64dd6b https://medium.com/@valyala/high-cardinality-tsdb-benchmarks...
- PaulWaldman 6y agoYes, I can refute that claim. TimescaleDB now provides built in compression. https://docs.timescale.com/latest/using-timescaledb/compression https://docs.timescale.com/latest/using-timescaledb/compress...
- mekster 6y agoYou didn't refute it technically. Are you saying the compression shrinks the data down to 2% on average? If the compression only makes the data 10 times smaller (I think I'm being generous with that ratio), it's still 5 times larger than the others.
- PaulWaldman 6y agoCompletely anecdotal, but I went from 50B to 2B per point using the compression feature. Mind you this is for slow moving sensor data collected at regular 1 second intervals. Prior to the compression feature, I had the same complaint. Timescale strongly advocated for using ZFS disk compression if compression was really required. Requiring ZFS disk compression wasn't feasible for me.
- douglasheriot 6y agoI've been using InfluxDB, but not satisifed with limited InfluxQL, or over-complicated Flux query languages. I love Postgres so TimescaleDB looks awesome. The main issue I've got is how to actually get data into TimescaleDB. We use telegraf right now, but the telegraf Postgres output pull request still hasn't been merged: https://github.com/influxdata/telegraf/pull/3428 https://github.com/influxdata/telegraf/pull/3428 Any progress on this?
- akulkarni 6y agoThere is a telegraf binary available here that connects to TimescaleDB: https://docs.timescale.com/latest/tutorials/telegraf-output-plugin https://docs.timescale.com/latest/tutorials/telegraf-output-... If you are looking to migrate data, then you might also want to explore this tool: https://www.outfluxdata.com/ https://www.outfluxdata.com/
- valyala 6y agoI believe PromQL [1] and MetricsQL [2] are much better suited for typical queries over time series data than SQL, Flux or InfluxQL. [1] https://medium.com/@valyala/promql-tutorial-for-beginners-9ab455142085 https://medium.com/@valyala/promql-tutorial-for-beginners-9a... [2] https://victoriametrics.github.io/MetricsQL.html https://victoriametrics.github.io/MetricsQL.html
- a10c 6y agoOne of the biggest quirks that I had bumped up against with TimescaleDB is that it's backed by a relational database. We are a company that ingests around 90M datapoints per minute across an engineering org of around 4,000 developers. How do we scale a timeseries solution that requires an upfront schema to be defined? What if a developer wants to add a new dimension to their metrics, would that require us to perform an online table migration? Does using JSONB as a field type allow for all the desirable properties that a first-class column would?
- valyala 6y ago90M datapoints per minute means 90M/60=1.5M datapoints per second. Such amounts of data may be easily handled by specialized time series databases even in a single-node setup [1]. > What if a developer wants to add a new dimension to their metrics, would that require us to perform an online table migration? Specialized time series databases usually don't need defining any schema upfront - just ingest metrics with new dimensions (labels) whenever you wish. I'm unsure whether this works with TimescaleDB. [1] https://victoriametrics.github.io/CaseStudies.html https://victoriametrics.github.io/CaseStudies.html