8 ms·
The current state of OpenTelemetry
- baby_souffle 3y agoDepends on who you ask. I am glad that the observability sector has standardized on a common protocol but my god are the reference implementations lacking.
- hinkley 3y agoPart of it is the spec. Stop letting Java people “help”. With friends like these who needs enemies. - ex Java person
- politelemon 3y ago> my god are the reference implementations lacking. Can you share some of your experience, what do you mean by that? Are there edge cases causing problems, or major missing features? Easy or difficult to use?
- abeppu 3y agoAs an example, Exemplars are part of the metrics spec [1]. The official python library says metrics status is 'stable' [2]. But there's an approximately 2-year old issue with no work on it, titled 'Metrics: Add support for exemplars', where the latest update is that no work has begun [3]. Nothing at a top-level of the opentelemetry-python project indicates that the project does not implement everything in the metrics spec, so if you wanted to use that capability, you are apt to discover it relatively late. [1] https://opentelemetry.io/docs/specs/otel/metrics/data-model/#exemplars https://opentelemetry.io/docs/specs/otel/metrics/data-model/... [2] https://github.com/open-telemetry/opentelemetry-python https://github.com/open-telemetry/opentelemetry-python [3] https://github.com/open-telemetry/opentelemetry-python/issues/2407 https://github.com/open-telemetry/opentelemetry-python/issue...
- sph 3y agoIn the Elixir library, we don't even have metrics. I went into the OTel rabbit hole for two days trying to understand how is it better to Prometheus, just to learn it doesn't even do the basic thing, just traces. I've mentally decided to just go Prometheus and ignore OpenTelemetry for the foreseeable future. It's one of those things big players are hyping to preemptively lock you in their solution, but it's actually just alpha-quality new tech and "boring" "old" tech like Prometheus or statsd are simply more functional and better supported in the wild.
- filmor 3y agoMetrics are implemented in the `opentelemetry_experimental` application. Last time I tried them, they were still a bit buggy but working (not complete, thiugh).
- notpublic 3y agoElixir opentelemetry works quite well with Tempo. Tempo does the metric generation [1] and writes it to Prometheus. Tempo also does Service Graphs which works great with context propagation [2]. Btw, metric generation is not enabled in Tempo by default. # tempo.yml overrides: defaults: metrics_generator: processors: [service-graphs, span-metrics] # Prometheus --web.enable-remote-write-receiver # Grafana.yml [feature_toggles] enable = tempoSearch tempoBackendSearch traceToMetrics [1] https://grafana.com/docs/tempo/latest/metrics-generator/ https://grafana.com/docs/tempo/latest/metrics-generator/ [2] https://hexdocs.pm/opentelemetry_process_propagator/OpentelemetryProcessPropagator.html https://hexdocs.pm/opentelemetry_process_propagator/Opentele...
- arwineap 3y agootel logging is completely missing from golang for example
- bbkane 3y agoI agree this should be there, but I also think in most cases, logs can be completely replaced by otel tracing - see https://www.infoq.com/presentations/event-tracing-monitoring/ https://www.infoq.com/presentations/event-tracing-monitoring...
- arwineap 3y agoThe example presented seems to also log, they just annotated the logs with span data > What we ended up implementing was a little tee inside the o11y library. As well as sending events to Honeycomb, we also converted them to JSON, and wrote them to stdout. That way, after sending to stdout, we then pumped off to our standard log aggregation system. This way, we've got a fallback. If Honeycomb is not working, we can just see our logs normally. We could also send these off to S3 or some other long term storage system if we wanted. I'd like to go a step further, and say that in addition to being worried about honeycomb being down, sometimes you just want to check with kubectl to get an idea what is going on. Our current projects are very log light because of the heavy tracing instrumentation, but it'd be nice to integrate this with the otel paradigms as they were originally intended
- eep_social 3y agotrace spans are structured logs, they just also happen to correlate into a tree, ymmv
- hinkley 3y agoThe silent failures by default.. I hate it. So. Much. Flames. Flames! On the sides of my face. Breath… Heaving breaths.
- Thaxll 3y agoI remember looking at the Go implementation, it did not look like Go code but it looked like someone was doing Java/C#.
- baby_souffle 3y ago> Can you share some of your experience, what do you mean by that? Are there edge cases causing problems, or major missing features? Easy or difficult to use? Just the general problem you get with big, slow moving OSS projects like this. Mostly just docs not current and a massive delta between certain languages; a feature is `stable` for some languages but not others which makes it hard to push for consistent otel roll out in a mixed-language environment. Some other "misc" points: - Google how to do $thing and you might find the proposed spec which gives example code ... that isn't what actually got implemented. That's a different link further down on your google results. - Python auto-instrumentation is ... fragile at best. It's not super clear if instrumentation is supported only with well known frameworks or just ... in general. I'd sure love some docs that explain how it works, too. - certain things require the collector use GRPC, others work with grpc or http... and I only found this out after googling an obscure error and reading through a _very_ long GH issue thread.
- pranay01 3y agoAgree, there being an open standard for instrumentation is a big win. Lots of work still needs to be done on showing more examples and making it more accessible to users & implementors. One other key area is resources which can help get engineers/implementors to get organizational buy-in
- neonsunset 3y agoC# has pretty nice integration with OTEL out of box (ASP.NET Core and otherwise, distributed as separate packages) https://learn.microsoft.com/en-us/dotnet/core/diagnostics/observability-with-otel https://learn.microsoft.com/en-us/dotnet/core/diagnostics/ob...
- yodon 3y agoCoincidentally used in for real for the first time today. Saved me an insane amount of time in how easy it made finding the root case of a perf issue.
- cowgoesmoo 3y agoMeh, I work in metrics observability and there's very little support for otel. Most new open source products are still based on Prometheus, which has much better SDK support than otel. I think it's a mistake for Otel to do its own thing instead of just building on top of Prometheus.
- MuffinFlavored 3y agoWhere it gets confusing: https://grafana.com/docs/grafana-cloud/send-data/otlp/send-data-otlp/ https://grafana.com/docs/grafana-cloud/send-data/otlp/send-d...
- wdb 3y agoPrometheus supports the OTLP format
- hinkley 3y agoBut OTLP is still hot garbage right now. If you send otlp to Prometheus it might get there, or it might all end up being dropped by a parsing error, because otelcollector is dumber than a box of hammers that have been through a rock tumbler.
- hinkley 3y agoI don’t agree with the communication patterns of either Prometheus or OpenTelemetry, but I’ll pick Prometheus next time I have to do telemetry. Unless there’s some fork of StatsD with tags that makes a resurgence.
- phillipcarter 3y agoAs a maintainer and end-user, my answer to this is...yes and no. It's important to clarify that, stability - something mentioned in the article - has several major definitions: - Stability in the specification - Stability in semantic conventions - Stability in the protocol representation - Stability in SDKs that can generate data - Stability in the Collector that can receive, process, and export that data Unfortunately, for many people, they may interpret "stable" in one of those categories as "stable for everything", and then get really annoyed when they find their language doesn't actually have stable support (or any support!) for that concept. What I'm most proud of in 2023 is all of the little things we made progress on with components that engineers have to materially deal with. On the website, we documented what feels like a million little things and clarified tons of concepts that people told us were confusing. Across all the SDKs, we fixed tons of little bugs, added more and more instrumentations, and completed the unsexy work to make metrics generation stable across most of our 11+ languages. The Collector added oodles and oodles of support for different data sources, and OTTL went from a neat component to a rock-solid general-purpose data transformation tool. There's so much more work to do, but I'm really happy about the progress.
- shipit1999 3y agoOpenTelemetry is a great concept, but in my experience not quite there yet. Docs especially fall into the common trap of handling the happy path hello world quickstarts, then become increasingly useless as you want to get beyond that to real life use cases. Given the inherent tradeoff of complexity that comes from trying to unify different approaches around one standard, sometimes it seems like things that should be simple are more difficult than they should be. I'm sure it will keep improving.
- mason55 3y ago> Docs especially fall into the common trap of handling the happy path hello world quickstarts, then become increasingly useless as you want to get beyond that to real life use cases. Yeah, Java is what I'm most familiar with. The "Getting Started" shows how to do some basic manual instrumentation and collect the output with curl. Then the "Next Steps" are just random things with no guidance about why I would or wouldn't choose any of them for my next step. But, ok, I choose "Automatic Instrumentation", that sounds promising. And it actually is really easy to set up auto instrumentation. But then at the end it says > After you have automatic instrumentation configured for your app or service, you might want to annotate selected methods or add manual instrumentation to collect custom telemetry data. Uh... no... after I have automatic instrumentation enabled I want to do something with the output The two major flaws in the docs seem to be 1. The common failure of docs to explain to users why they might choose one thing or another. "If you want to do x.. If you want to do y.." what if I don't know? 2. Because otel is agnostic to the consumer of the output, there's very little in the way of explaining how to get value out of what otel produces. To connect the dots, you really need to use the docs of your observability tool. Which I understand, but then most of them have their own setup directions because they want some extra fields included in the data, or they have their own fork, so not everything in the otel docs is actually usable. I'm not sure what the answer is. It's not like I expect otel to document how to build a dashboard in Grafana. And a lot of frustration I've experienced has been with the observability tools themselves. But at the same time, I always feel like the otel docs just don't get you anywhere close to getting value out of the library. Which is a shame, because turning on auto-instrumentation and seeing all your traces with literally no extra work is a magical moment.
- reindeerer 3y agoGive me something that isn't based on protobufs at wire / request level. CBOR with CDDL for a fully standards based approach that can work at any size of the stack
- pranay01 3y agoWhat's the issue with protobufs?
- reindeerer 3y agoThe first issue is that protobufs arent a standard. That inherently limits anything built on top of them to not be a standard either, and that limits their applicability Also depending on the environment you run in, can code size bloat vs alternatives can matter
- tonyarkles 3y ago> Aren’t a standard You mean like an IETF standard? That is true, although the specification is quite simple to implement. It is certainly a de-facto standard, even if it hasn’t been standardized by the IETF or IEEE or ANSI or ECMA. > inherently limits anything built on top of them to not be a standard either I’m not sure that strictly follows. https://datatracker.ietf.org/doc/html/rfc9232 https://datatracker.ietf.org/doc/html/rfc9232 for example directly references the protobuf spec at https://protobuf.dev/ https://protobuf.dev/ and includes protobufs as a valid encoding. > depending on the environment I’ve had several projects that ran on wimpy Cortex M0 processors and printf() has generally taken more code space in flash than NanoPB. This is generally with the same device doing both encoding and decoding. If you’re only encoding, the amount of code required to encode a given structure into a PB is very close to trivial. If I recall it can also be done in a streaming fashion so you don’t even need a RAM buffer necessarily to handle the encoded output. Do I love protobufs? Not really. There’s often some issue with protoc when running it in a new environment. The APIs sometimes bother me, especially the callback structure in NanoPB. But it’s been a workhorse for probably 15 years now and as a straightforward TLV encoding it works pretty darned well.
- deleted 3y ago[deleted]
- silveraxe93 3y agoThat's a terrible plot. I have no idea what the x-axis or the circle areas mean.
- hinkley 3y agoI thought I understood it the first time, and was looking again to explain it to you. Yeah I got nothing. Log-log plot? Why.
- bilalq 3y agoThe biggest issue with OpenTelemetry is how aggressively it's being pushed, despite not being mature enough. The AWS X-Ray team frequently suggests switching to OTel on bug reports and feature requests, but the performance and resource overhead of OTel collector for Lambda is just awful right now. It doesn't make sense for any performance sensitive workload. Beyond that, it gives off an "over-engineered" vibe. It's probably not, and the complexity of being a unified standard that can work across so many different variations is inherently going to need a lot of abstractions, but it feels so much more difficult to go through OpenTelemetry compared to an opinionated observability SaaS.
- codexb 3y agoI agree that it gives off an over-engineered vibe. I think part of that is that a lot of the "Getting Started" docs don't give you a feel for how to use the framework. It's more along the lines of "install this package and create this esoteric config file and this very particular telemetry logging will work". That's not a particularly useful walkthrough unless I have that exact same use case and want nothing more.
- onlyrealcuzzo 3y ago> performance and resource overhead of OTel collector for Lambda is just awful right now Presumably this means it's costly, which would be a reason for them to recommend it.
- hinkley 3y agoI resent that you are right.
- bilalq 3y agoI mean, sure, you can improve performance a bit by increasing the RAM/compute capacity on the Lambda. But it always adds a pretty steep overhead right now, no matter how much capacity you throw at it. https://github.com/open-telemetry/opentelemetry-lambda/issues/263 https://github.com/open-telemetry/opentelemetry-lambda/issue... https://github.com/aws-observability/aws-otel-lambda/issues/228 https://github.com/aws-observability/aws-otel-lambda/issues/...
- CornCobs 3y agoAs a relative outsider to the observability space, I have always wondered this: Is observability/telemetry only about engineering-related issues (performance, downtimes, bottlenecks etc.) or does it include the "phone-home" type of telemetry (user usage statistics, user journeys)? Looking through the websites of most of the observability SaaSes it seems to only talk about the first. Then how do people solve the second? Is it with manual logging to the server from the client?
- rafamct 3y agoIt is sometimes the second. Apollo (the GraphQL one) uses OpenTelemetry for tracing and monitoring reasons but also for usage tracking. When was a field last used, what frequency is it included in queries, etc.
- mongrelion 3y agoI believe right now this type of telemetry data is for whitebox monitoring for backend components
- renlo 3y agoI think usage statistics tend to require more retention time to discover user behavior and understand how to optimize revenue. In the general case people probably won't care much whether their widget was running at X% CPU on Dec 5th 2019 but they might care more about what percent of users did Y action on that date. When I worked on an observability team (not as an expert but as a general swe) we had two metrics pipelines; one was strictly usage statistics which came from the client, the other was purely server metrics but a subset of them were considered usage metrics which were aggregated and sent down with the client metrics for the folks upstairs.
- wdb 3y agoThey are working on RUM support in OpenTelemetry
- hinkley 3y agoI would think anyone trying to put an OTEL source on an embedded device (IoT) was out of his goddamned mind. OTEL assumes data sources have ample hardware and particularly memory. It periodically summarizes all traffic since startup instead of streaming things as they happen.