4 ms·
What always puzzles me about OpenTelemetry is that tracing, metrics and logs are all designed independently. I wish there was a way I could just annotate my cod
by EdSchouten 1mo ago
What always puzzles me about OpenTelemetry is that tracing, metrics and logs are all designed independently. I wish there was a way I could just annotate my code base once, and let the ultimate decision to expose something as a metric/log/trace be dynamic at runtime.
For example, if I look at a graph in monitoring dashboard and see something suspicious, I’d like to say: “The next time something like this occurs again, please save me a trace.” I should be able to just do that with a single mouse click.
I remember them releasing the tracing spec/SDKs and saying “now let’s move on to metrics/logs.” That never sat right with me.
- thorian1828i03 1mo agoOne strategy do to do that is to trace everything by default and select what to sample later, e.g. https://grafana.com/docs/grafana-cloud/observe-and-act/adaptive-telemetry/adaptive-traces/ https://grafana.com/docs/grafana-cloud/observe-and-act/adapt...
- EdSchouten 1mo agoIf I understand that correctly, it means your app always creates traces, and Grafana Cloud is responsible for sampling/aggregating. That may be prohibitively expensive in terms of CPU/network load. What I’m suggesting is that your apps by default only send metrics to your monitoring system, but that the monitoring system can specifically ask to “upgrade” metrics to traces. Or to log entries. The same thing with metric cardinality: by default, only report metrics in a fully aggregated manner. But do tell the monitoring system how they can potentially be broken up if needed (i.e., which labels to add).
- thorian1828i03 1mo ago> The same thing with metric cardinality: by default, only report metrics in a fully aggregated manner. But do tell the monitoring system how they can potentially be broken up if needed (i.e., which labels to add). How does the monitoring system have any of the context to add labels? That would only exist in application memory. Grafana went the other way - your app exports all labels, and then you selectively aggregate on ingest: https://grafana.com/docs/grafana-cloud/observe-and-act/adaptive-telemetry/adaptive-metrics/ https://grafana.com/docs/grafana-cloud/observe-and-act/adapt... > That may be prohibitively expensive in terms of CPU/network load. In practice I've not experienced this even on quite high request rates. While it isn't free, exporting everything has been cheap enough that the real cost in dollars spent is basically marginal (it's _storing_ the data that's expensive)
- EdSchouten 1mo ago> How does the monitoring system have any of the context to add labels? That would only exist in application memory. Indeed. If you have a protocol that doesn’t allow exposing that kind of information, then that only lives in application memory. But my suggestion is that it’s exposed.
- ffsm8 1mo agoYou're pitching a solution that's incredible brittle and unnecessarily complicated if you think about it in technical terms. For your feature to work you need bi-directional communication between the otel receiver and your application - that's still doable in general, but now you want a synchronous "upgrade" to traces. Now we're talking about a massive performance impact - and you need to somehow cache all otel data locally so they're available for the upgrade and only then submit then. It is a architecture that's not very smart, honestly. And precisely the reason why you'd simply submit everything and let the receiver figure out which samples it wants to keep - as thorian pointed out earlier.
- ragall 1mo ago> If I understand that correctly, it means your app always creates traces Yes, because otherwise what you propose requires modifying the binary in-place and that's too big of a security hole for lots of (production) environments. Some variants of that could work with an out-of-process method like Dtrace or eBPF, but that means mutating the kernel, even more of a no-no.
- PunchyHamster 1mo agoIt is very easy way to have your tracing infrastructure cost more than actual infrastructure.
- veqq 1mo agoYou can do that in Lisp, since you can arbitrarily redefine the wrapper to have such or other logic etc.
- MathMonkeyMan 1mo agoTracing is the most general of them, and the most expensive unless you're careful with the implementation. Trace spans are time-delimited units of "stuff that happened", with a tree relationship among the spans, and each span can have arbitrary tags (key/value pairs) and events (time/value). From that, if you chose, you could derive metrics and logs. The trick is to start with tracing and to actually put it in your program, rather than trying to mostly-automatically tack it on later.
- fuzzy2 1mo agoI just don’t get this sentiment. How would you represent metrics as traces? You cannot. Even reconstructing traces from logs would be challenging at best. How would you get, say, Garbage Collector metrics from logs or traces? You cannot. There is no magic bullet. Observability isn’t something you can just slap on and call it a day. While traces and logs might share superficial similarities, they are not the same. And metrics are something else altogether. Trying to somehow unify them would be a prime example of "wrong abstraction". > “The next time something like this occurs again, please save me a trace.” The building blocks for this exist. The observability platform must simply (haha) implement the pattern detectors and use them for sampling decisions.
- anygivnthursday 1mo agoI am not sure if this is what they mean, but e.g. with Micrometer in Java you can instrument your code once with observations that produces observation events, then you can register handlers that can turn them into metrics, or logs, or traces without having to instrument your code three times. https://docs.micrometer.io/micrometer/reference/observation.html https://docs.micrometer.io/micrometer/reference/observation....
- Flamkuchlo 1mo agoThe problem is not the instrumentation but the way everyone of them work. A metric is a point in time. A metric is very small but you have a lot of them. A log is when something is happening but you need to log it out. A logline is heavy and has a lot of context. User id, message, etc. A trace needs to start at the request level and tracing until the response. This is the slowest and heaviest operation. How do you decide when to suddenly do the trace and send it? IF you always do the trace, you have to pay for the overhead of that tracing constantly.
- spockz 1mo agoTechnically, you can use the same places in the code where you stop/start/fork traces to also be the places where you increment the counters/gauges, etc. Which I think the GP was alluding to when describing the micrometer solution. Similarly, you can derive metrics for log lines without having to emit the actual log lines. Then separately you can have log levels or verbosity levels that control to which level you actually emit traces/logs and/or roll up metrics.
- hobofan 1mo agoI don't think OTEL is necessarily "at fault" here. It's a split that's carried all throughout the observability ecosystem. e.g. in the Grafana suite of solutions you have Loki (logs), Tempo (tracing) and Mimir (metrics) to cover storage & querying for all three axis, as all of them have very distinct processing & performance characteristics. While it may intuitively may look like there is a large overlap in the three areas there is suprisingly little, and for the few parts there are (e.g. trace <-> log correlation), OTEL does offer a standard.
- spockz 1mo agoI think it is almost a inevitability where otel came as a standardised aggregate of OpenTracing (which was the same but only for tracing over multiple tracing implementations), logging, and metrics into a single observability standard without alienating all the individual supporting vendors. Historically, logging and metrics have been different problem domains with different implementations for ages. Now to your point: Note that tracing does get the most of love, and that it does include constructs to add logging and metrics into these traces (spans actually). So you could argue that they are trying to develop a single interface. > “The next time something like this occurs again, please save me a trace.” Well, if you want this you either need to propagate this predicate to all points that might be involved, or always emit all traces and have the predicate included in the filter. And then you need to be able to dynamically propagate this predicate from the system/ui where you click to where you filter. This is one of the reasons why we always propagate and emit traces and just post filter it in processing before it lands in the persistence layer.