8 ms·
I’m increasingly of the opinion that we should stop using “logging” libraries, and instead we should produce telemetry via events using something like OpenTelem
by glenjamin 4y ago
I’m increasingly of the opinion that we should stop using “logging” libraries, and instead we should produce telemetry via events using something like OpenTelemetry’s trace API.
These traces can be rendered as strings for people who like their logs in files
- rcarmo 4y agoMany tracing libraries get their data off logging libraries. Logging doesn't always mean "whoa, let's stop everything and write to a file", it is often a matter of buffering data or increasing counters until data can be sent out in various fashions. Log files are just one of those.
- glenjamin 4y agoThe key difference is my experience is not about destination - backends are as flexible as you say One key difference is having an event name + properties, rather than a human-readable description and maybe some fields if you’re lucky Another key difference is wrapping units of work rather than just scattering log lines in places you think need them somewhat arbitrarily
- rcarmo 4y agoMost decent logging frameworks allow for structured output. That change came about roughly when NewRelic became "the thing" to use.
- jayd16 4y ago>One key difference is having an event name + properties [...] Another key difference is wrapping units of work [...] Modern "enterprise" logging libraries already work this way, don't they?
- markmark 4y agoYou probably need to look into modern logging if you think that's how logging still works.
- tptacek 4y agoThe otel trace API is nowhere near ergonomic enough to replace logging. I think people mostly tolerate otel because the idiomatic way to use it is just a funcall and a defer at the top of a subset of your functions, and because if you're otel-instrumenting, you're probably already passing a context. https://pkg.go.dev/go.opentelemetry.io/otel/trace https://pkg.go.dev/go.opentelemetry.io/otel/trace
- henvic 4y agoI've played with trying to create an idiomatic wrapper for their Go module precisely because of this reason, but then didn't find time to continue with this. https://github.com/henvic/tel https://github.com/henvic/tel
- phillipcarter 4y agoFWIW tracing in Go is primarily made more difficult by having explicit context propagation and no helper to auto-close a span based on scope. If you use Python or .NET for example, it's implicit and you can "just create some spans" like you would a structured log somewhere in your app, then use built-in language constructs to have it handle closing for you.
- tptacek 4y agoSure. If it was a different language, or a different otel library, my opinion would probably be different. I do not think it's realistic to replace logging with otel in Go, though.
- phillipcarter 4y agoI think it's an inclusive OR here. When logging stabilizes in the Go SDK (which will be a while, perhaps a year) the idea will be that you can keep all your existing logging and then start to add tracing judiciously so they're easily correlated with requests. I think the juice is worth that particular squeeze for more folks at least. From there I would personally advocate for adding more manual traces rather than more logs, but it's understandable if it's a heavier lift to pass context around all the time. The more options the better.
- chrsig 4y agoI'm inclined to agree. Traditional logging has a number of problems. Developers tend to just stick `log.Info` (or whatever equivalent) anywhere. Using an http server as an example: The result is that one http request may produce dozens of log messages. But because multiple http requests are running at the same time, they all get interleaved in the actual log. They can be tied together after the fact if you're using a request id, but that requires post processing. It can be difficult to know when a request ends, so that means there needs to be an explicit "end of request" type message, otherwise assume that it could still be active until the end of the file...unless your log is rotating, in which case it could be in the next file, or the next. On top of that, it's a big performance hit. On every log message, you're serializing the message, acquiring a mutex, flushing it out to disk. Depending on the library, some of that may or may not get evaluated depending on message severity. It's a non-trivial overhead, and individual messages can cause quite a bit of contention for that mutex. Using something more like tracing, a log message gets attached to a context for the task, and it's not until the task is complete does anything get logged. This makes for a single log message that has a list of all the individual informational messages, as well as other metadata about the request. One mutex acquisition, one serialization, one file flush. It's at that point where tracing and logging are arguably the same thing.
- KronisLV 4y ago> It's at that point where tracing and logging are arguably the same thing. On some level I agree with this, but honestly tracing doesn't have nearly the same adoption that logging has (at least in my experience). Why? Because generally it's easier to approach logging, even if in ways that don't necessarily scale that well: like getting your application writing logs to a file on the server directory. Note: I mention Java here, but a lot of it also applies to Node/Python/Go/Ruby etc. Eventually they can be moved over to a networked share, have Logrotate running against them, or perhaps have proper log shipping in place. That is perhaps the hardest aspect, since for tracing (outside of something like JavaMelody which is integrated inside of your app) you'll also need to set up some sort of a platform for the logs to be shipped to and processed in any number of ways. Personally, I've found that something like Apache Skywalking is what many who want to look into it should consider, a setup utilizing which consists approximately of the following: - some sort of a database for it (ElasticSearch recommended, apparently PostgreSQL/MariaDB/MySQL viable) - the server application which will process data (can run as an OCI container) - the web application if you need an interface of this sort (can run as an OCI container) - an agent for your agent of choice, for example a set of .jar files for Java which can be setup with -javaagent - optionally, some JS for integrating with your web application (if that's what you're developing) Technically, you can also use Skywalking for log aggregation, but personally the setup isn't as great and their log view UI is a bit awkward (e.g. it's not easy to preview all logs for a particular service/instance in a file-like view), see the demo: https://skywalking.apache.org/ https://skywalking.apache.org/ For logs in particular, Graylog feels reasonably sane, since it has a similarly "manageable" amount of components, for a configuration example see: https://docs.graylog.org/docs/docker#settings https://docs.graylog.org/docs/docker#settings Contrast that to some of the more popular solutions out there, like Sentry, which gets way more complicated really quickly: https://github.com/getsentry/self-hosted/blob/master/docker-compose.yml https://github.com/getsentry/self-hosted/blob/master/docker-... For most of the people who have to deal with self-hosted setups where you might benefit from something like tracing or log shipping, actually getting the platform up and running will be an uphill battle, especially if not everyone sees the value in setting something like this up, or setting aside enough resources for it. Sometimes people will be more okay with having no idea why a system goes down randomly, rather than administering something like this constantly and learning new approaches, instead of just rotating a bunch of files. For others, there are no such worries, because they can open their wallets (without worrying about certain regulations and where their data can be stored, hopefully) and have some cloud provider give them a workable solution, so they just need to integrate their apps with some agent for shipping the information. For others yet, throwing the requirement over to some other team who's supposed to provide such platform components for them is also a possibility.
- akira2501 4y agoI'm of the opinion that we should use both. They're both fundamentally useful and solve separate problems, and I don't think there's any resource constraints that would make this an issue. A logging API that just lets me write either a textual string, a structured telemetry object, or just both would be ideal. Then I can have either source separately, or I can view how they generated events collectively.
- phillipcarter 4y agoOpenTelemetry's design says: why not both! Logging is still early days, but the idea is that you can take your existing structured logs and have them automatically correlated with traces that you can add later. (there's other use cases some can envision, but I'm team "tracing is better than logging", while recognizing that most apps don't start from scratch and need to bring their logs along for the ride)
- jen20 4y agoI'm in the position of working on a green-field system currently, and spent a bit of time looking at OpenTelemetry both as a logging system and for metrics and tracing. If OpenTelemetry can indeed suit that use case today, it's not evident from any of the documentation or examples, and the naive "go get" approach in a test project led to incompatible versions of unstable libraries being pulled in. Further, the OpenTelemetry Go library pulls in all kind of dependencies which are not acceptable: a YAML parser which doesn't even use compatible tagging and versioning via Testify, logr and friends, and go-cmp. Combined with the insistence of GRPC on pulling glog and a bunch of other random stuff like appengine libraries, you end up with a huge dependency tree before you've even started writing code. The result is I'm sticking with zerolog, and a custom metrics library, and will let Splunk do it's thing. It's a shame, as distributed tracing is very desirable, but not desirable enough to ignore libraries that have chosen concrete dependencies vs interfaces and adapters. Fundamental things like this need to be part of the standard library, or have zero external dependencies. Rust does (in my opinion) a better job here despite requiring on average more libraries to achieve things, because features can be switched off via Cargo.toml, reducing the overall dependency set.
- cube2222 4y agoI agree. I'm not necessarily a fan of OpenTelemetry, as every time I looked at it it was still "in progress" (though the last time was almost 2 years ago), but in general I agree. We started with tracing+logging, but as time passed we moved more and more stuff into tracing. But not just span-per-service kind of tracing - detailed orchestration on the level of "interesting" functions, with important info added as tags. Right now we use almost exclusively tracing (and metrics ofc) and it's working great, even if each trace often has hundreds of spans. In our case we use Datadog, which works well for most bits and purposes, but we also send all traces over Firehose to S3 so that we can query them with Athena at 100% retention. This way you get your cake and can eat it too.
- phillipcarter 4y agoIf you've got the time for instrumenting something relatively small, it's definitely worth picking up OTel again to see how it works for you. Tracing is stable in most languages (Go included) and there's a lot more support in each language's ecosystem. FWIW I think OTel will always be in a state of "in progress" though. There's more scenarios and signal types that will get picked up (profiling, RUM) and more things to spread across the ecosystem (k8s operator improvements, collector processors for all kinds of things, etc.) and more integration into relevant tools and frameworks (.NET does this today, maybe flask or spring could do it tomorrow). But I wouldn't let that stop you from trying it out again to see if you can get some value out of it!
- jayd16 4y agoI guess Golang is catching up but in something like C# or Java, the logging libraries already feel like tracing libraries. They already have contexts and structure. What are you proposing should change? Log files should default to a verbose binary representation that captures more? What do you think is not captured by today's logging?
- phillipcarter 4y agoTracing gives you correlation between "log lines" and across process/network boundaries by propagating context for you. You could replicate everything that tracing does by manually propagating around context-like metadata yourself, stitching the log lines together into a "trace". But why not just use an SDK that does that for you?
- jbjbjbjb 4y agoIt is a config away in C# world, and even when it wasn’t so in built it was very simple to add to structured logging.
- phillipcarter 4y agoIn the C# world, tracing is also just a config away. It's baked into the runtime. But you don't need to manually stitch together logs into a "trace" - it just does it for you.
- jbjbjbjb 4y agoWhat I meant was you can easily add a correlation id to your logs, and send those to wherever you collect and query the logs. Then using that you can easily see the log across your services. So there isn’t any manual stitching in the logging either.
- phillipcarter 4y agoYou can! And then if you want to establish causality from log to log, representing which part of a system calls which other part, you can add another id that lets you know which log is a “parent” of another. And maybe later, you may need to propagate metadata across a request, so you build that system and use it to better enrich your logs. And maybe later, you’ll need to correlate subsets of logs with other subsets of logs, so you build a way to “link” related sets of logs together. And so on, and so on. All of this is possible with “just logs” today and if it’s your jam, go for it. But these are reasons why tracing libraries exist.
- jyounker 4y agoOr you include the trace ID as one field in the structured log. That ID can be threaded through the log library.
- azlev 4y agoTelemetry and logs are distinct. It makes sense to gather metrics and it makes sense to catch sporadic messages of what's going on in details.
- arinlen 4y ago> I’m increasingly of the opinion that we should stop using “logging” libraries, and instead we should produce telemetry via events using something like OpenTelemetry’s trace API. That really depends on what you mean by "events". Even OpenTelemetry supports logging events. Logging is not replaced by emitting metrics, as they have very distinct responsibilities. > These traces can be rendered as strings for people who like their logs in files You're missing the whole point of logging stuff. The point of logging things is not to have stuff show up in files. Logging events are the epitome of observability, and metrics and tracing events end up specializations of logging events which only cover very precise and specific usecase.
- glenjamin 4y ago> Logging events are the epitome of observability, and metrics and tracing events end up specializations of logging events which only cover very precise and specific usecase. This is a common way of thinking about things, but having now worked with event-based tracing for the last few years I no longer believe traditional logs are necessary. If you do “good” logging you likely have entirely structured data and no dynamic portions of the main log message. This is almost exactly equivalent to a span as represented by tracing implementations. I have come to believe that these events are the core piece of telemetry, and from them you can derive metrics, spans, traces and even normal-looking logs.
- arinlen 4y ago> This is a common way of thinking about things, but having now worked with event-based tracing for the last few years I no longer believe traditional logs are necessary. You are free to use stuff as you see fit, even if you intentionally miss out on features. Logging is a very basic feature whose usecases are not covered by metrics events, or even traces. Logs are used to help developers troubleshoot problems by giving them an open format to output data that's relevant to monitor specific aspects of a system's behavior. It makes no sense to emit a metrics event with tons of stuff as dimensions and/or annotations that's emitted in a single line of code that's executed by a single code path when all you need is a single log message. Also, stack traces are emitted as logging events, and metrics events are pretty useless at tracking that sort of info. But hey, to each its own. If you feel that emitting a counter or a timer gives you enough info to check why your code caught an exception and what made it crash, more power to you. The whole world seems to disagree, given standard logging frameworks even added support for logging stack traces as structured logs as a basic feature.
- zgiber 4y agoLog the things that require a human to fix something every time they happen (Include useful info that’s needed for fixing). Create metrics of the rest: things that show patterns. Trace a percentile of actions that span over multiple components. I believe all the three: logging, metrics, tracing have their own uses, one can’t replace the other.
- jahewson 4y agoNice idea but I don’t want to have to take the hit of tracing every request.
- javier2 4y agoI dont disagree, but who the hell has time to setup and maintain the OpenTelemetry instrumentation and its myriad of other components.
- phillipcarter 4y agoWould love to learn which areas of OTel you find are onerous to set up and maintain. Experiences can differ from language to language which is why I ask. For example, with JS/Node there's a convenient SDK package and autoinstrumentation metapackage, but with Go you need to install a lot more stuff to get an equivalent experience.
- javier2 4y agoWe mainly use go and java, with different dbs and way too many http clients. But the instrumentation is even just a small part of it, we actually had most services fully instrumented with OpenTracing, but then they decided to abandon OpenTracing for OTel... And even then, the more difficult thing is running the trace ingestion without it crashing or running out of resources all the time, and balance that with just having everything disabled. Our most stable setup was just running with the in memory store so we at least have a trace for a few hours when we need it, but we essentially just gave up.