3 ms·
What are you using for container/pod logging?
by ktamura 10y ago
What are you using for container/pod logging?
- obeattie 10y agoGood question. We have a couple of approaches to this: * Every request that comes into our system is assigned a unique ID, which is propagated on every downstream call and returned in a response header. When logs are emitted during request processing, they are tagged with this ID. A system we've built in-house indexes these log events against their trace ID in Cassandra (on a separate cluster). This lets us take a failing, slow, or otherwise interesting request and look up all the things that happened to it during processing. Events in this system are TTL'd according to their severity – so an event at critical or error severity is kept longer than one at debug severity. * stdout/stderr from all our containers is forwarded to journald on each host. Logstash then pushes all these logs to Elastic (and also to permanent cold storage). This is useful to look at the "big picture" and means we can analyse all the logs in aggregate and makes it very obvious when something is _very_ wrong and causing a lot of requests to fail, but is less useful than slog for pinpointing a specific issue. It's worth also noting (since I find often a whole load of things can be mixed into logging) that we do not drive our monitoring off these logs, at least at the moment. We have separate systems for that.
- sandGorgon 10y agothis is awesome! how do you do this? a new header with a unique id... generated by something like lua+nginx. but then how do you pass this request from one service to another?
- amatix 10y agoFor us, a `X-Request-ID` header is generated by any app if it doesn't receive it from upstream -- but normally nginx or the CDN will generate it. There's a few nginx modules to do it, we use https://github.com/newobj/nginx-x-rid-header https://github.com/newobj/nginx-x-rid-header Most languages/logging frameworks have some sort of per-thread context (eg. Filters in Python, MDC in log4j, etc) to be able to tag log messages with. If you're using postgresql, you can call `SET application_name='{requestID}';` and that can be output as part of logs too.
- foxylion 10y agoWe do the same in our setup. A Apache HTTPd assignes a unique id (mod_unique_id) as a http request and response header. So any downstream system will get the request header and can attach it to the logs. (In our case we write json log and one field is the request id)
- gdubya 10y agoThere are quite a few monitoring products being built to solve this problem. Many of them are based around Zipkin (http://zipkin.io/ http://zipkin.io/)
- caniszczyk 10y agoI strongly suggest people look at the OpenTracing work too: http://opentracing.io/ http://opentracing.io/
- lobster_johnson 10y agoHow good is Cassandra at log-like data? Also, why the split between Cassandra and Logstash? Why not a single solution?
- obeattie 10y agoBecause of its disk layout, Cassandra is truly excellent at time-series, append-heavy data. In our setup, the data is partitioned by time bucket and n number of labels (one of which is the request ID). We may unify the two at some point, but there's no immediate need to do so. While the write use-case is quite similar across both, the read use-case is quite different: slog requires reasonably low latency reads soon after the data is written, data can age out after 2-30 days depending on severity, and sometimes dropping events is acceptable. It would be acceptable for reads from the "archival" system to take minutes or even hours, the data should be kept forever (or for a long time), and dropping events is never acceptable.