3 ms·
While you do need to create tables for your stats, you are able to ALTER TABLE just as you can for normal Postgres without problem. It may not be as pain-free a
by RobAtticus 10y ago
While you do need to create tables for your stats, you are able to ALTER TABLE just as you can for normal Postgres without problem. It may not be as pain-free as NoSQL in that regard, but building tooling for this or automating this is certainly possible.
Alternatively if you do want pure blog storage, Postgres's JSONB datatype works as well, although with some performance trade-offs. This can work quite well if you have some structured data that lives alongside other unstructured data. [1][2]
Building or putting a REST front-end in front of PostgreSQL should not really be an issue (we built one for our hosted version), and there are already a fair number of PostgreSQL clients/libs. That said, for ease of use we are already thinking of adding an HTTP interface.
For Grafana: We were actually working on our own connector to support PostgreSQL, but found out that one is already in the works by the Grafana team (which will work out of the box with Timescale, because each of our nodes look like PostgreSQL to the outside world). This doesn't seem like it'll be an issue for very long.
We've mentioned retention policies elsewhere, but you do bring up a good point. Instead of dropping data, support for aggregating data in older chunks to a more coarse-grain resolution is something we are already looking into.
[1] https://www.citusdata.com/blog/2016/07/14/choosing-nosql-hstore-json-jsonb/ https://www.citusdata.com/blog/2016/07/14/choosing-nosql-hst...
[2] https://blog.heapanalytics.com/when-to-avoid-jsonb-in-a-postgresql-schema/ https://blog.heapanalytics.com/when-to-avoid-jsonb-in-a-post...
- koffiezet 10y agoI think you're a bit missing my point. My view is maybe a bit limited to the 'ops' side of things, but creating and maintaining tables completely defeats how I'm currently using metrics, and I suspect this applies to most people using them when I look at the available ops-targeted metrics collector tools: Collectd, Telegraf, Intel's Snap, ... Going over the list of their available modules/plugins/sources should give you an idea of how realistic it is to maintain tables for all those metrics if you were planning on adding support to them. The main reason time-series databases are 'hot' these days is because ops jumped on them, and I think understanding how metrics are being used there and for what purpose is the key here. My first reaction when something generates 'events' or data - whatever that might be - is simply to push them into a time-series database without thinking about them. Is it data from snmp (a switch, firewall, ...), generic machine stats, database stats, application metrics, ping statistics, time until an ssl certificate expires, ... you name it - I don't care, maybe it'll be useful, maybe not. I don't even dare to estimate the amount of different metrics I'm currently tracking. Real-world example: did I think I had to know what the sizes of my different caches on my ZFS storage units were? Not at all, but Telegraf pushed them anyway - so whatever. They actually ended up being very useful tracking slow-downs on our fileshares caused by some rogue process scanning all files on them, completely trashing the caches in the process. I had the data right there at my disposal, and I didn't really knew the details of what exactly I was tracking until I had to take a closer look. These use-cases are something TimescaleDB's approach on it's own would be completely unsuitable for. As you mentioned, JSONB has performance trade-offs and blindly using that for everything sounds like a recipe for disaster. It also only addresses one problem, complexity to use is another, and everyone having to define their own structure and insert queries is yet another: there is no standard. While there are quite a few time-series out-there, most popular tools support the most important-ones, but I don't see how you can add support for TimescaleDB since the data-structure is completely undefined. That's why I mentioned the tgres (which I had encountered before). Something like that, backed by a Postgres database with TimescaleDB's extension could be very interesting. Doing the 'generic' metrics through a simplified, well defined interface with some fixed table-structures, while still allowing way more powerful 'direct' metrics to address niche needs sounds very interesting to me. Currently however, I only see TimescaleDB useful within one well-defined application, where it can be very valuable. In the grand scheme of 'ops' things however, it feels too limited and inflexible, and not really like a time-series database. > For Grafana: We were actually working on our own connector to support PostgreSQL, but found out that one is already in the works by the Grafana team Ah yeah, I forgot that the Postgres datasource type was added in Grafana, yes that would do. Other thing is time-series-specific functions, but I'm not really familiar with all stuff Postgres offers, I know it's a lot - but I couldn't find stuff like regression analysis/derivative functions. Or maybe it's my google-foo is failing. Providing metrics-specific functions in the Postgres extension should be possible though.
- JimNasby 10y agoSomething that might be very interesting would be combining https://www.torodb.com/ https://www.torodb.com/ with Timescale. ToroDB is database agnostic (though they prefer Postgres). If you extracted the appropriate keys (including the timestamp) out I expect the combination would be very powerful.