10 ms·
Zabbix, Time Series Data and TimescaleDB
- linsomniac 7y agoExperiences with Zabbix? I tried it back around a decade ago and wanted to like it, but didn't find it very reliable. And now the details are escaping me. I ended up sticking with Nagios and Opsview. Around 5 years ago I switched to a templated Icinga2 config and have been pretty happy with that, but it's pretty low level.
- skullborg 7y agoI’ve run Zabbix with thousands of monitored hosts. It’s not perfect, and it requires some bending to just how Zabbix wants things done, but it’s nice. We have it monitoring all manner of stuff, hardware, power, cooling, services, batteries, weather, network, disks, etc
- newaccoutnas 7y agoI've had a similar experience to you with Zabbix in the past. We have it bundled with some HPC stuff I support curently, it's ok but I prefer Sentry/TICK/Prometheus shaped things that we also run. If you're happy with Icinga2, stick with that. I've used that too at a previous gig and found it better that Zabbix, but my personal take on it. YMMV
- leonroy 7y agoZabbix is a bit opaque to tune and the support forums aren't super helpful unlike say Nagios. That said it does some really cool stuff like tree walking across all the HP switches on our network, auto monitoring all ports it finds and then reporting on their stats and on any UP/DOWN states for every port. Good for detecting unauthorized usage or a device which is rebooting itself. Its IPMI support is also pretty good, we had it monitoring Supermicro IPMI interfaces with zero issue. It handles vSphere and auto scans the entire cluster, adding all guests and monitoring them without needing to install an agent on every VM. All in all a very good solution with some very cool features, but a steep learning curve and not much help on their forums although the docs are pretty good.
- unixhero 7y agoYeah I agree. Zabbix sucked when I tried it many years ago. Definitely not going near it again.
- anorwell 7y agoOne thing I like about zabbix is the excellent grafana plugin, which provides a very good ability to view and ack host-by-host alerts from within grafana. That said, I'm not very familiar with the alternatives.
- Jnr 7y agoZabbix looks like shit and feels like it was made in 1995 but it is great once set up. If better visuals are needed, I would hook it up to Grafana. I have previously used Grafana with Graphite as backend but it was too unreliable. If it actually works with Zabbix then it could be the perfect match.
- bloopernova 7y agoMy experience with Zabbix has been positive. I deployed it 3 years ago and it's been solid since then. It monitors a few dozen CentOS VMs and a bunch of JBoss/JMS instances. One feature I particularly like is the zabbix_send command, which I use to push the status of shell-scripted Borg backup jobs into Zabbix.
- wbh1 7y agoSurprised to see Prometheus hasn't been mentioned yet, and even Nagios is being mentioned as a better alternative. My company (higher-ed, ~100k combined students/fac/staff) is desperately trying to get away from Nagios. Once you get Nagios to the scale where you have to implement mod_gearman, you've gone too far. I'd recommend taking a look at Prometheus[1]. It has its own _very_ performant TSDB, there's exporters for just about everything, it's the defacto way that things like Kubernetes expose metrics, and it has first class support in Grafana for visualization. We POC'd Zabbix, Icinga, ScienceLogic, Instana, Sensu, and Prometheus. Prometheus was our favorite. Take a look at the comparison between it and other popular monitoring products to see if it fits your needs though [2]. [1] https://github.com/prometheus/prometheus https://github.com/prometheus/prometheus [2] https://prometheus.io/docs/introduction/comparison/ https://prometheus.io/docs/introduction/comparison/
- tecleandor 7y agoThe problem I have with Prometheus is, I have most of my nodes in very closed networks I don't have control (Healthcare) and I can't set up proxies so Prometheus can reach them, I can only go outside. So, by now, my best option seems to be InfluxDB, which doesn't look bad to me.
- linsomniac 7y agoI've been using InfluxDB for ~3 years now for storing metrics (almost exclusively via Telegraf, a few custom ones), and it has been great! It replaced a collectd setup and dramatically decreased load across my fleet. When I first started using it, it was pretty early and had some issues. In fact, I nearly trashed it. I also didn't like the pull vs. push model from Prometheus. They ended up resolving the InfluxDB issues I was having right as I was about to give up on it, and it's been solid since. I use it with Grafana to generate graphs of system use. I set it up before TICK was a thing.
- h1d 7y agoI was about to like InfluxDB but ever since people say it eats memory and your data, I stopped caring. https://github.com/VictoriaMetrics/VictoriaMetrics/wiki/FAQ https://github.com/VictoriaMetrics/VictoriaMetrics/wiki/FAQ ("How does VictoriaMetrics compare to InfluxDB?")
- eeeeeeeeeeeee 7y agoAbsolutely terrible. Like you said, totally unreliable. Scaling it is insanely difficult. Documentation is weak. Their APIs seem like an afterthought and performance was pretty bad. Nagios is not great, but it’s reliable and when it breaks you can figure it out.
- lima 7y agoTimescaleDB confuses me. Postgres is an OLTP database and their disk storage format is uncompressed and not particularly effective. By clever sharding, you can work around the performance issues somewhat but it'll never be as efficient as an OLAP column store like ClickHouse or MemSQL: - Timestamps and metric values compress very nicely using delta-of-delta encoding. - Compression dramatically improves scan performance. - Aligning data by columns means much faster aggregation. A typical time series query does min/max/avg aggregations by timestamp. You can load data straight from disk into memory, use SSE/AVX instructions and only the small subset of data you aggregate on will have to be read from disk. So what's the use case for TimescaleDB? Complex queries that OLAP databases can't handle? Small amounts of metrics where storage cost is irrelevant, but PostgreSQL compatibility matters? Storing time series data in TimescaleDB takes at least 10x (if not more) space compared to, say, ClickHouse or the Prometheus TSDB.
- qaq 7y agoThere are a ton of projects that will never outgrow TimescaleDB so if you have in house PostgreSQL expertise looks like very decent option.
- claytonjy 7y agoI can't say for sure, but shouldn't insertions be way quicker in Timescale, because the index-changes are limited to the most-recent subtable only, and it's still row-based? We're considering a move from OpenTSDB to Timescale currently, and something that stands out in Timescale is the wide-table format; we get bundles of metrics at each tick, and having them aligned makes usage easier, and perhaps also saved us some space over having the timestamps repeated per metric.
- valyala 7y agoConsider moving to time series database with PromQL support. It is much easier to write typical queries over time series data in PromQL than in SQL or Flux. See https://medium.com/@valyala/promql-tutorial-for-beginners-9ab455142085 https://medium.com/@valyala/promql-tutorial-for-beginners-9a...
- 7y ago
- sreeramb93 7y agoMy experience with timescaledb is - it does not support gorilla encoding. So the storage needs for it is very high.
- SEJeff 7y agoGorilla TSDB format paper for those who might not get the reference: https://www.vldb.org/pvldb/vol8/p1816-teller.pdf https://www.vldb.org/pvldb/vol8/p1816-teller.pdf
- techntoke 7y agoZabbix is like a step up from Nagios. I don't know how they can even stay relevant with Prometheus.
- marknadal 7y agoThe biggest problem I've had with timescale systems is managing the SSDs/HDDs underneath. Having to resize/grow/stripe/etc. them is a pain. So we came up with a clever solution that batches chunks to S3: https://www.youtube.com/watch?v=x_WqBuEA7s8 https://www.youtube.com/watch?v=x_WqBuEA7s8 $10/day for 100M records (100GB data), all costs! And best yet, reduced DevOps! Very practical, super simple.
- deleted 7y ago[deleted]
- enordstr 7y agoTimescale engineer here. Just want to point out that you can also attach additional disks using tablespaces, which are fully supported on hypertables. With a few simple commands, this allows you to add new disks and move old disks out of rotation while still being able to query the old data on them.
- valyala 7y agoWe prefer Google Cloud durable persistent storage. It may be started from a few GBs and then resized online up to 64TB per instance. This allows saving money by resizing disks only when needed. Such disks cost $40/TB/month. See https://cloud.google.com/compute/docs/disks/add-persistent-disk https://cloud.google.com/compute/docs/disks/add-persistent-d...
- boomskats 7y agoEvery time TimescaleDB is brought up, I feel the need to point people to their shadily worded proprietary licence[0], and pg_partman[1]. Do the same benchmarks against a pg_partman managed partitioned db and you'll get the exact same performance. We do, at least - 150k or so metrics per second, 10 columns per metric. Not trying to crap on the TimescaleDB guys, I've found a lot of their writeups extremely useful and can totally see how their commercially supported product fits. However, I like to see pg_partman at least mentioned somewhere in the article/comments. It's awesome and does the same job. [0]https://github.com/timescale/timescaledb/blob/master/LICENSE https://github.com/timescale/timescaledb/blob/master/LICENSE [1]https://github.com/pgpartman/pg_partman https://github.com/pgpartman/pg_partman
- detaro 7y agoLink to a file that actually contains said license: https://github.com/timescale/timescaledb/blob/master/tsl/LICENSE-TIMESCALE https://github.com/timescale/timescaledb/blob/master/tsl/LIC...
- cevian 7y agoClarification: this license is only for Community and Enterprise Features that live in the `tsl` subdirectory. The vast majority of our code (everything not under `tsl`) is licensed under Apache 2.
- deleted 7y ago[deleted]
- mfreed 7y ago(Timescale cofounder here) Hey, just wanted to clarify: the vast majority of TimescaleDB code is Apache2, and you can easily compile (and we ourselves build & distribute) Apache2-only binaries. When we announced a new license in December, we didn't relicense any code, we just said that some future features will be available under a Community or Enterprise License. The code under this "Timescale License" is clearly marked and in a separate subdirectory, and for virtually all users (except the public cloud DBaaS providers), the community features are free. This is the actual top-level LICENSE file in the repo: https://github.com/timescale/timescaledb/blob/master/LICENSE https://github.com/timescale/timescaledb/blob/master/LICENSE And here's a blog post discussing in more depth: https://blog.timescale.com/how-we-are-building-an-open-source-business-a7701516a480/ https://blog.timescale.com/how-we-are-building-an-open-sourc...