3 ms·
InfluxData is arguably playing catch-up with Thanos, Cortex, and other scale-out Prometheus backends for the metrics use case. Given that, I wonder why they dec
by sciurus 6y ago
InfluxData is arguably playing catch-up with Thanos, Cortex, and other scale-out Prometheus backends for the metrics use case. Given that, I wonder why they decided to write a new storage backend from scratch instead of building on the work Thano and Cortex have done. Those two competing projects are successfully sharing a lot of code that allows all data to be stored in object storage like S3.
https://grafana.com/blog/2020/07/29/how-blocks-storage-in-cortex-reduces-operational-complexity-for-running-prometheus-at-massive-scale/ https://grafana.com/blog/2020/07/29/how-blocks-storage-in-co...
- pauldix 6y agoThose two systems are designed to work with Prometheus style metrics, which are very specific. You have metrics, labels, float values, and millisecond epochs. I'm not totally sure how they index things, but I would guess that it's by time series with an inverted index style mapping for the metric and label data to underlying time series. This means they'll have the same problems with working with high cardinality data that I outlined in the blog post. InfluxDB aims to hit a broader audience over just metrics. We think the table model is great, particularly for event time series, which we want to be best in class for. A columnar database is better suited for analytics queries, and given the right structure is every bit as good for metrics queries.
- Thaxll 6y agoAlso: https://github.com/VictoriaMetrics/VictoriaMetrics https://github.com/VictoriaMetrics/VictoriaMetrics from the guy that did FastHTTP ( the fastest HTTP lib for Go ).
- throwawaytsdb 6y agoI agree, they appear to be playing catch up on many fronts. Notably, with Cortex, Tempo, and Loki, Grafana Labs seem to have pulled way ahead in advancing a successful open-source cloud observability strategy. InfluxData have a long history of writing (and rewriting) their own storage engines, so choosing to do it again is unsurprising. I guess this sort of hints that the current TSM/TSI have probably reached their performance and scalability limits and will be EOL before too long. What I find interesting is that this project is already almost a year old and only has six contributors (two of whom look like external contractors). It seems more like a fun side project than the future core of the database that is supposed to be deployed into production next year.
- nemothekid 6y ago>What I find interesting is that this project is already almost a year old and only has six contributors I don't think this is a fair criticism. Postgres only has 7 core contributors - and influxDB is far less complex targeting a much simpler use case.
- throwawaytsdb 6y agoPostgreSQL might have only seven people on the "core team", but the active contributors list is much longer: https://wiki.postgresql.org/wiki/Committers https://wiki.postgresql.org/wiki/Committers Furthermore, hundreds of people have been responsible for its development over the years: https://www.postgresql.org/community/contributors/ https://www.postgresql.org/community/contributors/ As an additional, possibly more relevant example - the Cortex project has 158 contributors: https://github.com/cortexproject/cortex https://github.com/cortexproject/cortex My larger point is that, in contrast with the new project, InfluxDB currently has ~400 contributors. I'm certain that many dozens of those were involved in getting the current storage engine to a stable place. And now that hard work is on a path to being deprecated by moving to a completely new language and set of underlying technologies. Taking the project from a handful of contributors to a production-ready technology within an existing ecosystem is a non-trivial task. I'm sure it will come together eventually, but the commitment to ship it "early next year" seems unlikely to me.
- pauldix 6y agoWe'll be producing builds early next year. Those won't be anything we're recommending for production. Our goal is to have an early alpha in our own cloud environment by the end of Q1. I stress alpha. But we'll also have a bunch of tooling around it (which we've already built for the other parts of our cloud product) to backup data out of band, monitor it, shadow test it against production workloads, etc. We're also building on top of the work of a bunch of others that have built Arrow and libraries within the Rust ecosystem. When the open source GAs, I don't really know. But we're doing this out in the open so people can see, comment, and maybe even contribute. Who knows, maybe after a few years you'll be a convert ;)