3 ms·
Splunk and Sumologic both have query languages, and store the raw log lines now, and let you parse and analyze them later. ELK doesn't have anything like that;
by dfboyd 9y ago
Splunk and Sumologic both have query languages, and store the raw log lines now, and let you parse and analyze them later.
ELK doesn't have anything like that; you have to configure ELK to understand the logs before it can index and store them, and it has no query language allowing you to re-parse and analyze them afterward -- all you can do is search on the fields you've already indexed. ELK can't, for instance, store 20G of logs today with no parsing applied, and let you parse out key-value pairs and graph all the latency numbers next week.
If you have more money than time, use Splunk.
If you are on AWS, dump your logs into Redshift -- you already know SQL and you don't have to learn another query language; and it's easy to see how to run extract/transform/load jobs.
If you're on GCP, I'd investigate BigQuery before bothering with Splunk or Sumologic.
- scapecast 9y agoIf you are a start-up, I second the approach of dumping logs into Redshift. Lots of start-ups are doing exactly that (including yours truly). Few things to keep in mind when you do that: Create separate schemas for your raw data and your analysis. Dump the logs into a "raw schema", run your aggregations, and write the results to a "data schema". Once your data grows and you're running more reports, it will make your life much easier. Separate your users from the start. Create a user each for your data loads, for your aggregations and for your ad-hoce analysis. As you ramp up query volume, the separation of concerns will make it easier to use the Redshift workload manager and keep concurrency high. Ping me if you have more questions about the set-up. lars @ intermix dot io
- aprdm 9y ago> ELK can't, for instance, store 20G of logs today with no parsing applied, and let you parse out key-value pairs and graph all the latency numbers next week. This isn't entirely true, you can have a generic logstash-* index and dump log lines into it (say sending from syslog). All you need is a timestamp and a string. Later on you can reindex that data at anytime, using logstash (and/or elastic curator) and then reinterpret that data into a different index in which you will be able to graph on Kibana.