6 ms·
Author here! Really excited to release this at KubeCon; happy to answer any questions you might have.
by netingle 8y ago
Author here! Really excited to release this at KubeCon; happy to answer any questions you might have.
- zellyn 8y agoI'm curious about this, because at Square we maintain our own homegrown log aggregation system, and it's not really a core competency. While a lot of our logging needs seem like they would be fulfilled by this system — because we attach trace IDs to log messages, and because (at least in Payments) you can usually find the appropriate trace ID by searching for a Payment ID, which could be annotated too — there are definitely many times I've copy/pasted the text in quotes from a log-generating line of code in a Java or Go file, to find out if it's being executed, or as a handle into a subsection of code/logging. In the linked design doc, you include a motivating tweet near the top, saying, “just give me log files and grep, I am dying”. But unless I'm misreading things, there's no `grep` here. Right? I'm guessing you could narrow down (using metadata) and then grep, but if the narrowest metadata you have is app name and time range, you're still going to be grepping over a lot of data…
- netingle 8y agoIt’s defiantly something that’s missing from the readme, and perhaps not that obvious in the grafana explore view either - but it is there! You can push a regexp match server side and have that distributed to each Loki node, giving you distributed grep. Will make it more obvious. Davkals has an iteration of the UI that makes it a separate field, which will also help.
- danlimerick 8y agoThere is some documentation about the Loki search syntax in the Grafana docs: http://docs.grafana.org/features/explore/#logs-integration-loki-specific-features http://docs.grafana.org/features/explore/#logs-integration-l...
- samstave 8y agoPlease make it more obvious with exact examples of how to do this.
- GordonS 8y agoMaybe I'm not understanding this - the docs say Loki is all about storing compressed log data with metadata, such that only the metadata is indexed. Are you saying you can search the compressed, unindexed data using regex? If so, wouldn't that potentially be incredibly slow?
- gouthamve 8y agoThe good thing about this is that the grepping can be parallelised and distributed on to several nodes. Having said that, once you select the relevant metadata right, you should be able to narrow it down enough for the queries to be snappy enough. While this will definitely be slower than something that indexes the contents, you'll be able to store much more in Loki at much lower costs.
- GordonS 8y agoYeah, I am thinking about the worst case here, but never underestimate the power of your users to perform very silly queries! For your hosted service, will you put in place any restrictions on, for example, the size of the time range that can be queried? Also for your hosted service, will the degree of parallelisation vary by pricing tier?
- zacmps 8y agoWhat are you using to run the regex? ripgrep could make up for some of the loss from not having it indexed.
- ecnahc515 8y agohttps://github.com/grafana/loki/blob/master/pkg/iter/iterator.go#L252:6 https://github.com/grafana/loki/blob/master/pkg/iter/iterato... Looks like the Go regex lib, which isn't super performant, so it could potentially be improved if it ends up being an issue.
- programd 8y agoI was going to suggest you look at OKlog [0] but it seems that the project is recently archived and no longer maintained. A pity because I really like its simple and minimalistic design. It does still work of course and the code is all there. Loki would do well to look to it for some inspiration. I particularly like the super simple install, and lack if the need for some clunky UI to tail and grep your distributed logs. [0] https://github.com/oklog/oklog https://github.com/oklog/oklog
- netingle 8y agoAgreed - Loki was heavily influenced by OKLog. We really like its ease of use and simplicity - hopefully we managed to get some of that right with Loki. OTOH we felt like we needed a little bit of metadata and index to help find the right logs - the brute force approach of only having grep doesn't seem like the right tradeoff to me. We use a minimal label index to narrow down the search space + then let you do distributed grep... There is a lightweight CLI for Loki too, you don't have to use the Grafana UI: https://github.com/grafana/loki/blob/master/docs/logcli.md https://github.com/grafana/loki/blob/master/docs/logcli.md
- sandstrom 8y agoThat's great. With a lightweight CLI you can run this in development and just use the CLI. Then add Grafana to your production environment.
- zellyn 8y agoThanks! Always love Peter's stuff. We actually have a perfectly viable working system internally. We're just (always) considering alternatives :-)
- bauerd 8y agoWhat if I have a service that logs quite verbosely and shows anomalous behaviour over say 10 minutes? I assume I'd have to scan all those log messages sequentially as they're not indexed at all, right? Will Grafana offer a UI for doing so or would I have to dump the log segment and grep it?
- netingle 8y agoLoki is design to compliment Prometheus; as such we envisage you using Prometheus metrics to isolate the service and time range exhibiting the anomalous behaviour (by looking at latency and error metrics, for instance) and then "pivot" to Loki to see that logs. As Loki uses the same label metadata as Prometheus, that pivot is automatic and almost "magic" - showing you the relevant logs for a given PromQL query. Grafana is offering a logging UI for Loki in the upcoming v6 release called Explore; you can enable it on the master builds right now, see https://grafana.com/blog/2018/09/21/grafanas-explore-ui-taking-a-deeper-dive-into-data-with-prometheus-queries/ https://grafana.com/blog/2018/09/21/grafanas-explore-ui-taki...). It makes is super easy to start exploring and sifting through your logs. Also, Loki allows you to push regexp matches server side, so you can distribute the "grep" among multiple machines for extra points :-)
- bauerd 8y agoThanks, makes sense. Seems like a great idea, kudos!
- shadycuz 8y agoGood luck today, will be watching from my office ;)
- GordonS 8y agoI clicked through and read this on the grafana page: > Loki is meant to be complementary to existing solutions like Elasticsearch and Splunk that do full text indexing Can you elaborate a bit on this please? If I'm already using Elasticsearch, Splunk or the like, why would I want to add on another, less powerful logging service? (not trying to be a dick, genuinely want to understand why I'd want this!)
- gouthamve 8y agoSimple answer: Cost. When debugging you'd want as much info as possible and you'd want to be able to simply tail + grep it. I've been told to log less because the amount I was logging would burn a hole in the pocket when deployed to production. Sometimes, at scale people only send WARN (maybe even only ERROR) and above in production which is sometimes not enough when trying to debug a system on fire. While ELK does a great job of indexing the contents, and if you depend on it for BI, you should definitely still keep it for use-cases where just select+grep won't suffice. But for just storing logs, and being able to select, stream and grep laaarge quantities of logs, Loki will come in handy.
- GordonS 8y agoSo for example, you might send only ERROR level to Splunk, and everything to Loki? Or maybe you'd send everything to Splunk, but with a small retention period, whereas you'd send everything to Loki but with a 90 day retention period (or whatever)?
- sandstrom 8y agoLooks awesome! Three questions: 1. It sounded like it was easy to set key-value metadata in the promtail tool. Things like hostname, availability zone, etc. But can you also append metadata via the log-files themselves? Basically, our log files are JSONL (http://jsonlines.org/ http://jsonlines.org/) and look like this: { "user-id" : "abc", "client-type" : "mobile", "etc" : "…", "messages" : ["Parsing incoming http request", "Saving user data, valid model", "Finishing up http request, sending response to client"] } { "user-id: " "def", "etc" : "you get the idea…" } { "third log line here" : "and so on" } Can we ingest these into Loki and have the "user-id" metadata appended as key-value labels to each message? 2. Sometimes it makes sense to group related log lines. In the example above, we'll have ~20-30 log lines from a single http request. It would be convenient if we somehow could group them together, for example based on a unique X-Request-Id value. And then use that in the UI to see all related log lines together easily. 3. We currently store metadata about log-lines that is numeric. For example, things like request time. Will it be possible to query on that type of numerical value, to e.g. find all the log lines where requests took more than 500 ms?
- wikibob 8y agoI think the problem with setting key-value metadata via the log output, is that permits unbounded cardinality. The current design inherently limits cardinality to the number of pods you have running and the various labels applied to them. For 3 I'd say that doesn't make sense with this design. I'd suggest taking a look at something like scalyr.com. You can configure a log parser and then your logs become both queryable, and you can create time-series on numeric fields such as request time and look at the 99th percentile, min, max, etc.
- grumpydba 8y agoHi! Is there a way to plug it with alertmanager and trigger alarms if a specific message is found in the monitored log ? Ie corruption message in a postgresql instance's log.
- GordonS 8y agoFew more questions: 1. Batching logs will allow better compression ratios and bigger blobs (which means lower per-operation costs), but must be balanced with the risk of data loss - what is your strategy here? 2. Will this handle multi-line logs? Say a regex matches part of a multi-line log, will it then return all the lines for that log? 3. Can you add your own labels, or are you limited only to those assigned by Loki? 4. If you are limited to labels assignd by Loki, how will you handle labelling as you expand out of k8s and accept logs for other sources (e.g. syslog)?
- GordonS 8y agoThought of another one - logs really have 2 timestamps associated with them: the time the log was generated, and the time the log was ingested by Loki. Will Loki be able to parse the log generation time out of logs (where it's included, and it usually is), or will it only use the log ingestion time for time range searches?