4 ms·
We do acknowledge that storage is not our strong suit, but we think the benefits of having SQL and being able to integrate easily with non-time-series data is a
by RobAtticus 10y ago
We do acknowledge that storage is not our strong suit, but we think the benefits of having SQL and being able to integrate easily with non-time-series data is a big win itself. Certainly if storage is an issue for you, TimescaleDB is probably not the right choice. But if it isn't, and you are doing more advanced queries (including JOINs or across multiple metrics), TimescaleDB provides a compelling solution in our view.
- endymi0n 10y agoIt's not just the storage itself (which is really cheap nowadays), but more of the fact that every byte read and written needs to go through the processor, pollutes RAM and poisons caches. Less storage usually translates into direct performance gains as well. Also, if you find yourself JOINing timeseries on more attributes than the timeline itself, you should question whether you really have a solid use case for a timeseries DB. That being said, always good to see competition in the market, especially if it's built on such a rock solid product and community. I really like that Postgres is "good enough" solution for almost every use case by now besides relational data, be it document storage, full text search, messaging - or time series now. Nothing wrong with having less stack to worry about, especially for prototypes and small scale projects!
- RobAtticus 10y agoAll good points. And especially agree about your last paragraph -- another benefit I didn't highlight is anyone familiar with Postgres does not have to learn a new part of the stack.
- cevian 10y agoOne more thing to add: Postgres index usage via index-only-scans go a long way to mitigating performance issues of wide rows (although, admittedly not disk-space issues). This allows good performance on columnar rollups.
- pgaddict 10y agoI think it's pretty clear Postgres is getting to a situation where the storage format becomes the limiting factor for a lot of workloads, particularly verious types of analytics. Timeseries are example of yet another type of data that would benefit from different type of storage. There already were some early patches to allow custom storage formats, chances are we might see something in one of the next Postgres versions (not in 10, though).
- crudbug 10y agoI totally agree with all the points here. Maybe Postgres needs pluggable storage engines [0]. I read Postgres 10 might have this, but looks like it will miss the deadline. [0] https://www.pgcon.org/2016/schedule/events/920.en.html https://www.pgcon.org/2016/schedule/events/920.en.html
- jeltz 10y agoPluggable storage will definitely miss PostgreSQL 10 since the feature freeze is later this week.
- anarazel 10y agoAnd there's not even a credible proof-of-concept patch.
- jnordwick 10y agoGood time series performance is more than just using column-based storage. You also need a query language to take advantage of this and the ordering guarantees it gives you. SQL while it has tried to reinvent itself, is a very poor language for querying TS databases.
- akulkarni 10y agoFrom personal experience, not sure I'd agree with that statement. SQL may be limiting for some time-series use cases, but for others it's quite rich and powerful. I won't pretend that SQL solves everyone's time-series problems, but we've found that it goes pretty far. That said, we may have to get a little creative to support some specific use cases (e.g., bitemporal modeling). Still TBD. Also, I agree that SQL isn't for everyone (the popularity of PromQL is evidence to that). But a lot of people have been using SQL for a while (personally, since 1999), and there is a rich ecosystem (clients, tools, visualization tools, etc) built around it.
- jnordwick 10y agoIt most definitely is.. LEAD and LAG are about all you get, and they are painfully slow. SQL was made to be order agnostic, and attempts to make it most order-aware don't quite work. A good time series database is build on table order and lets you exploit it. SQL is abysmal for any time series work. And temporal and bitemporal databases (despite the name) are orthogonal to the aggregation and windowing issues that make time series difficult in SQL or a row-oriented database. The are just a couple of timestamps and where clauses to support point-in-time queries. Maybe this is why so many time series databases fail. People making them often don't seem to fundamentally understand the issues, Very few, such as Kx and KDB, seem to understand them.
- anarazel 10y ago> It's not just the storage itself (which is really cheap nowadays), but more of the fact that every byte read and written needs to go through the processor, pollutes RAM and poisons caches. Less storage usually translates into direct performance gains as well. FWIW, I think that's currently a good chunk away from being a significant bottleneck in postgres' query processing. But we're chipping away at the bigger ones, so we'll probably get there at some point.
- johnnyhillbilly 10y agoYou are mentioning implementation-specific issues that may have more than one solution. If one goes to the heavy contenders in this space, e.g. Teradata, you may expect: - DMA for data retrieval - A suitable and linearly scalable network layer with Remote DMA - Row, columnar and hybrid storage options - Utilization of CPU vector options - Etc The analytical database has become a commodity. I really like Postgres, but I would still do a very careful analysis of my business needs if I were to choose a DBMS when there is such a strong range of commercial options available.