29 ms·
Why DNSFilter replaced InfluxDB with TimescaleDB
- LogicX 9y agoI'm Mike Schroll, CTO of DNSFilter - Happy to answer any questions about our experiences with InfluxDB or TimescaleDB over the last 2 years.
- camel_gopher 9y agoWhen you talked about 150M queries per day; are those inserts or reads? That's about 1,700 per second, which to be honest doesn't seem like a lot for time series metrics ingest. I would expect a single node on most TSDBs to be at least ~100 times that performance. Can you talk about the data you were ingesting? Was it numerics, text, or something else? (Disclosure, I work for TSDB provider IRONdb http://irondb.io http://irondb.io)
- LogicX 9y agoHi there -- We are currently ingesting 150M/day -- though it's not evenly distributed -- probably peaks around 3,000qps Agreed that it's not too much yet. I think the element which kills us is needing to do rollup tables, and then query against both the raw data and rollup data for our customer analytics dashboard. The data is a combination of date/time, strings, and ints. 22 fields.
- ryanworl 9y agoI know I mentioned this in another comment, but you really should check out Clickhouse. They have a table engine for exactly this purpose. You create a table with the raw logs, then a materialized view (or another actual table) which declaratively does the rollup for you in real time. https://clickhouse.yandex/docs/en/table_engines/aggregatingmergetree/ https://clickhouse.yandex/docs/en/table_engines/aggregatingm... A full example: https://www.altinity.com/blog/2017/7/10/clickhouse-aggregatefunctions-and-aggregatestate https://www.altinity.com/blog/2017/7/10/clickhouse-aggregate...
- pauldix 9y agoInflux founder and CTO here. I figure I should comment because it seems this post is a bit dated on information. That being said, I think this is one of those cases where based on where InfluxDB was at the time, they simply picked the wrong tool for the job. We were looking at use cases with hundreds of thousands of series. It just wasn't designed for what they were trying to do (for whatever version they were running at the time). We have much higher aspirations now, but I'll get to that in a bit when I talk to each point raised in the post. As I mentioned, it would be helpful to know which version of InfluxDB was tested against, but I'll try to cover each of the points. 1 - Cardinality. High cardinality has long been a known problem for InfluxDB. We started work on this almost 15 months ago to create a new feature called TSI or Time Series Index. It was significantly harder for us to develop than we thought it would be. It can be enabled on version 1.5, which was released yesterday. We'll have a detailed whitepaper on how it works and what techniques we used coming soon. It shows significantly lower memory usage for high cardinality write workloads. The other side of this is query, which I'll talk about in a bit. 2 - Rollup tables is a known issue for us. Continuous queries were originally supposed to do this, but the problem is that it falls over at scale. There are some other weird things about how it works where it won't pick up lagged data collection. It's a weak spot for us, Prometheus, OpenTSDB and a bunch of other solutions. From what I can tell, Timescale doesn't offer this feature either. See https://github.com/timescale/timescaledb/issues/350 https://github.com/timescale/timescaledb/issues/350. It sounds like DNSFilter implemented this themselves using their Kafka ingest pipeline and computing them and inserting into Timescale tables (which for now is exactly what we recommend with InfluxDB). If I'm guessing wrong, I'd love to hear the detail about how this works in your solution. 3 - Ingestion performance vs. Query performance. It sounds like query load causes ingestion performance to degrade. I really don't know without more information. In the post Mike said they were working with InfluxDB in 2015 and 2016 and didn't continue upgrading. So they could have been running 0.11.0, a version that is now over two years old. We've made significant improvements since then. However, I will say that our current query engine is a known weak spot. This is why we're investing heavily into IFQL (our new language and engine). The new language also addresses the InfluxQL doesn't actually operate like SQL, which is something he mentions at the top of the post. I agree that can cause frustration, which is why we're designing our new language around this use case. I think the functional paradigm makes more sense than SQL for time series. Slides from a recent talk I gave about the engine: Slides: https://speakerdeck.com/pauldix/ifql-and-the-future-of-influxdata https://speakerdeck.com/pauldix/ifql-and-the-future-of-influ... Video: https://www.youtube.com/watch?v=QuCIhTL2lQY&list=PLYt2jfZorkDqmNVloKpJGp-47KVeKcUPb https://www.youtube.com/watch?v=QuCIhTL2lQY&list=PLYt2jfZork... 4 - Resource utilization. It uses RAM, particularly on high cardinality workloads. Yes, we know this and have known it for a long time. Running tests on 1.5 with TSI enabled, the RAM utilization is significantly lower. However, TSI is only one side of the issue. The other side is querying high cardinality data. Our current query engine will eat many resources trying to do this. It's something we're addressing with the IFQL engine. You can actually use the new engine as a separate process against the InfluxDB 1.5 release. However, it's still under heavy work and we haven't begun the performance optimizations. I'm confused how Mike talks about about their query processor and then immediately dives into Kafka, which is only relevant in the ingestion pipeline. In fact, our recommended configuration is to route writes through Kafka before sending to InfluxDB if you're operating at scale. It's how we're designing 2.0 from day 1 and that work has already started. In an analytics pipeline, you want to separate your write pipeline from production query/storage. 1 - Ease of change: schema changes in InfluxDB aren't easy. Postgres supports alter table commands. How do these work on large tables? Has Timescale (or Postgres) solved the problem of these kinds of operations taking the DB down for a while? I think that's one of the strengths of having Kafka in their ingestion pipeline to protect against it (or other DB issues). We still need to do work here and I know it's a weak spot and it's honestly just a very hard problem (for how data is stored on disk for us). In the future, at the very least we'll enable users to kick a schema change off and have it run in the background while keeping the production infrastructure up. Depending on the change it might be a long running background task that requires rewriting many things. 2 - Performance - Impossible to address without knowing their data, their queries, and really testing it on a new version. We're optimizing all the time. I guarantee you that the query performance you see in InfluxDB today won't be anything close to as good as what it'll be a year from now. 3 - No missing data - I don't know of any issue in the current release (or many previous releases) of Influx that just drops writes. This might be something to do with Continuous Queries since it sounded like it was about rollups. It's a known issue with that feature and it's being completely reworked in 2.0 to work in a guaranteed way, at scale. It's a long running problem and one that is hard to get right. Again, it sounds like Mike and team implemented it themselves through their ingestion pipeline and then attributed the gain to Timescale. You can do the same thing with InfluxDB and it's our recommended architecture. We tell customers to do their downsampling I'm not sure what is going on with deletes. You can do deletes in Influx and you can delete by a time range, which would give you a specific point. That being said we've done significant work in the last three releases to improve how deletes work and particularly how they work for deleting large amounts of data. For loads, I'd really like to see a comparison against the 1.5 release with TSI on compared to Timescale. A release that's 1.5 to 2 years old with InfluxDB is completely different that what is current, as many people that have been watching the project over the last 4 years has seen. Finally, Mike does make a case at the end for Timescale just being SQL and Postgres so they found it easy to use because it's familiar. For better or worse, we're making the bet that SQL isn't the last only true API for working with time series data. Some developers will prefer SQL and some will prefer other approaches. For InfluxDB, we're making a bet that other approaches will be better for developer productivity.
- marknadal 9y agoDang, 3B records/day is very impressive - and I'm a competing distributed systems database engineer! Our tests were doing about 100M+ (100GB+) a day (this was 2 years ago), although on significantly smaller hardware, but I don't think it would've hit 3B even on the Intel Xeon you guys used. Stats: https://www.youtube.com/watch?v=x_WqBuEA7s8&t=1s https://www.youtube.com/watch?v=x_WqBuEA7s8&t=1s Just wanted to say, really impressive work! Respect.
- ryanworl 9y agoDid you evaluate Clickhouse?
- LogicX 9y agoI did not. One of the benefits from TimescaleDB is the ease of development and using standard queries and indexes to deal with the data. It would be another system for us to spin up on.
- pauldix 9y agoYeah, I'd think about that for the analytics use case. It's interesting technology and I always keep an eye out for what Influx can learn from these other projects. This CloudFlare post is on my readlist to learn more about it: https://blog.cloudflare.com/http-analytics-for-6m-requests-per-second-using-clickhouse/ https://blog.cloudflare.com/http-analytics-for-6m-requests-p...
- cevian 9y agoThe issue for an operational workload like this with Clickhouse might be the lack of direct support for UPDATE and DELETE. The workarounds required would add additional complexity, I think.
- jimaek 9y agoWe use influxdb as well at http://www.dnsperf.com http://www.dnsperf.com No issues so far except of course the initial difficulty correctly setting up the tables and continuous queries. The main problem for us is that clustering is available only to the enterprise version which is way too expensive for a small self funded company like ours. I wish they offered more options without any support and a big discount.
- pauldix 9y agoFor smaller users or cases we generally suggest InfluxCloud, our offering in AWS, which uses the clustering. Did you look into that?
- jimaek 9y agoYes but again it's too expensive to support our data. We rent servers with 16threads 128gb ram and 250ssd ssd raid for ~$100/month. The cloud option would cost thousands for the same resources. I understand that small customers like us are not interesting to bigger USA companies. But it's still something I wish was possible
- jxub 9y agoI wonder if they also considered KDB+ for that.
- dr_faustus 9y agoI find it fascinating how almost every No/NewSQL database in almost every use case niche (document, time series, key value, etc.) gets blown out of the water by Postgres either from day one or a couple of Postgres releases later (sometimes using a plugin). It just goes to show how great their technology is. Plus, its very mature, well understood, safe, secure and stable. And its much more likely to still be around (maintained) 10 years from now than the current NoSQL flavor of the month. After we had some pretty horrendous experiences with several NoSQL dbs (mostly MongoDB, CouchDB) on client projects, we strongly urge all clients to just use Postgres. The only exceptions are redis (for caching, queues, etc.) and elasticsearch (for ... search), which are just very convenient complements to Postgres. We have never found a single instance where the postgres performance was not as good or better than the NoSQL alternative.
- chatmasta 9y agoWhat about when you want to store documents with an unpredictable schema? That is, a collection where the field names and types could be different in every document. Granted, I've never actually encountered that situation.
- dogecoinbase 9y agoPostgres has exceptional JSON support.
- matthewmacleod 9y agoThe JSON support in Postgres is in my experience much more consistent than that of e.g. MongoDB, and I understand performance is better too. Postgres really is awesome :)
- gilbetron 9y agoDruid still crushes Postgres, so I'm happy with our decision. However, I get a lot of sad faces when eager engineers come to me and ask which "cool" database they should use to meet the needs of Project X. "Postgres, unless you can prove otherwise" is always my answer. It is a great piece of technology, but it isn't the answer to everything ... just most everything ;)
- wildchild 9y agoInfluxdb is a new fancy cloud bullshit to scam you. Don't fall for it. Eventually it will get rekt like everything rekts. You can't replace decades of hard work of postgresql engineers with a fancy toy.
- odiroot 9y agoI evaluated InfluxDB at my previous company. Settled for TimescaleDB in the end due to the querying power of Postgres. Influx had some real quirks with nested queries (wrong data being returned). TimescaleDB is probably a bit slower and less compact but with our data (a few GBs/day) it wasn't our biggest problem.