4 ms·
InfluxDB creator here, happy to answer questions and add more commentary here!
by pauldix 8y ago
InfluxDB creator here, happy to answer questions and add more commentary here!
- ezrast 8y agoI tried out InfluxDB a while ago in my spare time and was intrigued by the feature set, but ultimately couldn't get past the abstruse query language, especially coming from the simpler and more flexible PromQL (not being able to do ad-hoc math across time series was a big deal for my use case). I'm eagerly looking forward to giving it another shot with Flux and have super-high hopes. What does the data model for time series look like in 2.0? Mostly the same as 1.x, or has that gotten more flexible as well?
- pauldix 8y agoFor now we take writes in the 1.x line protocol so it’s still measurement, tags, and field. However, Flux doesn’t really make that a requirement. So in the future we plan on having a way to write series in without requiring a field or even a measurement. Once the planner gets the data to the Flux processing engine it views everything thing as a table of data with columns and records. So it’s much more flexible in how we can represent data.
- ezrast 8y agoSounds good; thanks for the reply!
- valyala 8y agoFYI, the article mentions that InfluxDB will eventually support PromQL: > With InfluxDB 2.0 we support both push and pull models out of the box. Eventually we’ll also support querying via PromQL out of the box
- zphds 8y agoI really dig the line protocol. Pretty simple. Any HA features? Sharding to look out in 2.0? Or is the general idea to set streaming relays of influxdb tsm's and treat HA as an L7 proxy routing problem (shadow metric traffic using envoy for example). How do people handle this in their production setups. Curious to know. It would be cool if the query engine could talk to multiple shards spanning multiple machines for dealing with high cardinality series.
- pauldix 8y agoRight now we’re prioritizing work on the single open source server and our cloud service, which has a very different design. Flux will be able to query multiple servers and combine their results (in OSS), but that would be a building block for some HA or clustering. So you could certainly layer in your own HA solution. We’re still working out what if any clustered for federated features will exist in open source.
- bithavoc 8y agothe fact the clustering is not part of the core is a show stopper for me
- wmf 8y agoAnd not making any money would be a show stopper for them. This is probably never going to change so complaining about it is probably pointless. (I'm not affiliated with Influx, but I strongly believe developers need to get paid with money.)
- s17tnet 8y agoI totally agree, I am a developer too but their offering is outrageously expansive.
- exabrial 8y agoTickScript is indeed the killer feature, the problem is debugging is difficult because of the language itself is esoteric. Simply adding the ability to put printf statements to the console would be a game changer. The ability to write tests outside of chronograf would be a game changer. The ability to mock inputs would be a game changer. The problem is you've re-implemented the wheel and now you have to build the debugging ecosystem that is a well beaten path for other languages. Love your guy's product. It's also hard to give you guys money, because there's no offering you make that quite fits one of our needs.
- pauldix 8y agoThanks for the TICKscript feedback. That’s all stuff that we want to have addressed in Flux. This alpha release doesn’t have that yet but printf, a test runner built into the influx CLI, and test inputs and outputs are all on the near term roadmap
- alexk 8y agoThat's my feedback on TICK as well, it's pretty hard to debug and implement something like this: https://github.com/gravitational/monitoring-app/blob/master/resources/alerts.yaml#L116 https://github.com/gravitational/monitoring-app/blob/master/... I looked at flux, and it seems pretty compatible with TICK script, although I don't have a clear understanding on how to make it easy to write alerts/queries and debug them yet. I would be curious to try out REPL: > We plan to provide a flux command line program that exposes a REPL and talks to various data sources. Especially interested to see how easy it would be to edit and troubleshoot a multi-line transformation query like this one in it: cpu = data // only get the last 5m of data |> range(start: -5m) // only get the "usage_user" data from the _measurement "cpu" |> filter(fn: (r) => r._measurement == "cpu" and r._field == "usage_user") Why create a new language vs using javascript or lua? e.g. your flux example above could have been just: let cpu = data.range({start: "-5m"}).filter(func(r){return r.measurement == "cpu" && r.field == "usage_user"}) Is there any specific feature of flux that requires a new language? I read your blog post here: https://www.influxdata.com/blog/why-were-building-flux-a-new-data-scripting-and-query-language/ https://www.influxdata.com/blog/why-were-building-flux-a-new... and it brings some valid points on Flux vs SQL, however I would be interested how would it compare as Flux vs Javascript or Flux vs Lua :)
- willvarfar 8y agoAre there any changes to the data storage level? Optimisations etc? And can data points be incremented instead of the current field-replacement crap when you get new points with the same tag set?
- e-dard 8y agoIn the last few months we have made quite a few improvements to data storage and indexing. Features that are available by default in 2.0, which are significant changes from releases earlier than, say, 1.6 include: - Significant TSM encoding and decoding performance improvements. - The TSI index will be on by default. - Queries that use the same tag key/value filters will be answered from the index more quickly using an LRU cache. - Field keys will now be indexed in 2.0, making filtering/grouping on field keys more efficient. - Improvements to how series are extracted from the index, and points data from the TSM engine, which helps with memory performance for queries. - Significant performance improvements to measurement deletion. > And can data points be incremented instead of the current field-replacement crap when you get new points with the same tag set? Can you elaborate on that?
- willvarfar 8y agoExcellent info on the engine improvements! I want to use influx to store _statistics_ not _events_. Basically, my data points are tag-sets and counts. There are several ways to achieve this; for example, you can send the events to influx and have continuous queries to gather the statistics. That doesn't work well when you have a lot of events, and where they arrive out of order and at high latency, etc. So what you typically end up having to build is a stats thing that sits in front of influx and tracks the counts of events with particular tag sets in particular time buckets, and then keep uploading these to influx. And there are two ways to do that: 1) you are not stateful and you keep uploading deltas and incrementing the nanoseconds to avoid data-point collision; you can then get the data out of influx with sum() on the fields and grouping by whatever the time bucket is. I tried this and influx grinds to a halt eventually. 2) you are stateful and track the totals outside influx, and keep uploading a newly-written data-point to overwrite the fields for that bucket in influx. This is much less data in influx and much easier to query, avoiding sum() etc. Its like I end up with something in front of influx doing what I want influx to do. What would greatly simplify life is if the line format which looks like this: measurement,tag1=x,tag2=y,tag3=z f1=total,f2=total timestamp could look like this: +measurement,tag1=x,tag2=y,tag3=z f1=delta,f2=delta timestamp and in the second case, where the line is prefixed with a + sign, influx knows to add the fields if the data-point collides with another rather than overwrite them. This would mean that people trying to store statistics in influx could add to those statistics statelessly. A massive simplification. I've had other problems, like I have way more than 1M series. Its painful. My influx boxes hit iowait far too often, which is weird because the boxes have more RAM than the total dataset.
- mattashii 8y ago> Our vision for 2.0 is to collapse the TICK Stack into one cohesive and consistent whole ... After InfluxDB 2.0, what applications would I have to run to be completely tick-compatible? I ask this because I would rather not have to run a DB on every server because Telegraf was fully merged into InfluxDB.
- e-dard 8y agoHi, just to clarify - Telegraf is still stand-alone. You will not need to run InfluxDB 2.0 on every host that you need Telegraf on.
- hunta2097 8y agoI love InfluxDB, apart from the need to have separate retention policies when reducing granularity over time. Are there any plans for a more unified method of performing continuous queries - so that we can query high granularity and older, downsampled data at the same time? That would be a killer feature for me.
- pauldix 8y agoAbsolutely. We won't address that in the initial release of 2.0, but there will be ways to get it done. The eventual solution will probably revolve around using the tasks system to downsample into buckets of different retention. Then using a function in Flux at query time to look at the metadata of buckets in the query and the time range and selecting the precision based on that. We should be able to show examples of how to do this in Flux later this year.
- hunta2097 8y agoAwesome, I'll look forward to it. Thanks Paul!