3 ms·
There's several, depending on use case. Each make different tradeoffs, so you have to decide what's important for you. For example Prometheus (which I work on)
by bbrazil 10y ago
There's several, depending on use case. Each make different tradeoffs, so you have to decide what's important for you.
For example Prometheus (which I work on) is great at reliable monitoring and powerful processing of metrics at high volumes, but it'd be unwise to use it for event logging or customer billing.
If you're doing IoT or event logging then InfluxDB might be a good choice for you, though if you're doing more text-based logging then Elasticsearch is nearer to what you're looking for.
https://docs.google.com/spreadsheets/d/1sMQe9oOKhMhIVw9WmuCEWdPtAoccJ4a-IuZv4fXDHxM/edit#gid=0 https://docs.google.com/spreadsheets/d/1sMQe9oOKhMhIVw9WmuCE... is one comparison of the various open source options.
- pwernersbach 10y agoIn addition, there is influx-mysql: https://github.com/philip-wernersbach/influx-mysql https://github.com/philip-wernersbach/influx-mysql , if you're looking for a MySQL compatible solution.
- lobster_johnson 10y agoWe're using Elasticsearch for event logging -- where "event" means analytics event, e.g. a page view -- and it's fantastic. The aggregation support is superb. We initially used Influx, but it could not perform well at the time (0.8). Our events are also heavily label-based. Basically, we do ETL at the time of write, collecting multiple documents into one mega-event, which is a complex, nested JSON document. It may have perhaps 150-200 fields. A single event may be something like "clicked button X". By storing the original document, we can aggregate based on any field value, including text and scalar fields, without having to think about a schema or about planning ahead of time what fields should be indexed or not. ES handles the rest pretty well. To do the same thing with Influx or Prometheus I suspect we'd have to reverse this and store the document as the labels, along with a single count (1) as the "metric". I don't know how well Influx etc. scale with number of unique label values, though I'd love to find out. The last time I read about this, I think they recommended not going overboard with them. What's different with business analytics is that the end product is typically multidimensional rollup reports over large time windows (number of page views per customer per web property per month, comparing by 2015 vs 2016, for example), and it's almost all "group by count", sometimes "count distinct" or averages. Whereas "rate per second"-type metrics aren't used anywhere in our app, for example.
- pixelmonkey 10y agoI also use Elasticsearch for time series site analytics use cases. I gave a talk about Elastic{on} about it last year. Apologies for the email gateway for the video, but you can also see my slides here: https://www.elastic.co/elasticon/conf/2016/sf/web-content-analytics-at-scale-with-parse-ly https://www.elastic.co/elasticon/conf/2016/sf/web-content-an... We found that as we scaled it up, we couldn't really keep the data in raw form, so we had to build rollup documents that cover 5-minute and 1-day buckets. Do you use the same trick, or is the number of pageview events for you manageable enough that you just keep it all raw?
- lobster_johnson 10y agoWe haven't reached that stage yet, fortunately; though at some point someone will want to do a big multi-year aggregation report across all indexes and still expect it to take not more than a few seconds. My ideal solution would be one that rotated the dataset into historical rollups on a daily basis, so that we only stored the raw data for today, and gradually merged earlier entries at lower granularities. However, I haven't thought much about how to do that with Elasticsearch. I can see a way of doing it by embedding the value in the field label, and using the field value as a count, but Elasticsearch really doesn't like lots of unique fields; you shouldn't be using more than a few hundred at most in a single installation (across all indexes).
- manigandham 10y agoAt that point, you can just use Apache Drill over raw JSON files: https://drill.apache.org/ https://drill.apache.org/