7 ms·
OTel isn’t going well
- hn_acker 2mo ago(TFA author is not me.)
- brikym 2mo agoI've never found instrumentation to be a huge issue. Sure it takes more effort but you get a lot more value once you understand _business_ events.
- jiggawatts 2mo agoThe alternative is vendor lockin, $$$, and spotty support for complex environments with zero chance of ever getting 100% coverage. At least with Open Telemetry, anyone can write an OTLP "source" using free, open specifications, and it'll "just work" with dozens of third-party "sinks". That's huge! Sure, there's a lot of experimental tags on semantic conventions, but at the end of the day, that's not that critical. It's just data: most sinks don't "interpret" these tags, they just display them as-is, so changes aren't breaking changes.
- GauntletWizard 2mo agoThe alternative is Prometheus (which is freaking great) and Jaegar (which is freaking great), each alone. This is better, because Otel is trying to put two distinct things (monitoring and metrics, distributed tracing) into one package, because they know how to use neither. Neither Prometheus metrics nor Jaeger traces are magic bullets. Neither of them are complicated, either, and in fact the fact that they're not complicated is their greatest strength. You can and should understand every facet of what they entail. You should build the (very small) shims that they need for your company's framework every time. It's not hard. It's not hard because it's not complicated. The fact that it's not complicated seems to break people's brains. They are accurate because they're simple and they're easy to work with because they're simple, and OTel is neither.
- firesteelrain 2mo agoI’ve built custom Prometheus metrics very easily and had node exporter pick up the .prom files. Python and bash scripts reading and translating. Node exporter runs on my Prometheus server next to Blackbox Exporter. Blackbox Exporter handles TLS expiry metrics.
- cyberax 2mo agoOTEL metrics are a bit awkward, but they work just fine with Prometheus. Jaeger uses the OTLP protocol nowadays. So it _is_ OTEL.
- arcanemachiner 2mo agoWhat, so people don't like OTel, but they like Jaeger, which implements an OTel spec? (I'm a noob to this subject, if that wasn't obvious.)
- deleted 2mo ago[deleted]
- cyberax 2mo agoYep. Kinda like people hating Obamacare but loving the ACA. Jaeger does not implement all the OTEL features, though. It's specifically focused on traces rather than metrics.
- GauntletWizard 2mo agoJaeger doesn't really implement an otel spec - Otel wrapped itself around Jaeger.
- nunez 2mo agoI believe you can still use Zipkin with Jaeger
- morganherlocker 2mo agoPrometheus is so easy to add and if you need more scale, there is mimir and a few other options with similar client semantics. I really can't imagine reaching for a framework APK that tries to anticipate every possible thing I would want telemtered, and is inevitably missing all the domain specific derived channels I need. Even prepackaged Prometheus exporters are usually overkill.
- bilalq 2mo agoOTel is so frustrating. If it wasn't shaping to be the clear winner in the space, I wouldn't complain about it as much. But today: 1. Every major vendor is still in some weird alpha/beta support for OTel even after all this time. 2. The performance hit is substantial and makes you question what the point of performance instrumentation is if you need twice as much compute/RAM to run the same workload now. 3. Serverless runtimes pay a heavy penalty for cold starts with OTel. 4. You're basically forced to run both gateway collectors and edge collectors for any realistic usage. 5. You still need to configure destination exporters in unique ways. This leaves you questioning what the value of OTel was. 6. Vendors that go beyond the scope of what OTel covers still need their own bespoke instrumentation. What was the point of any of this then?
- cyberax 2mo ago> 4. You're basically forced to run both gateway collectors and edge collectors for any realistic usage. You most certainly don't. You can run your app (especially if it's "serverless") without the collector agent. App-to-agent and agent-to-sink use the same protocol, so all you need to do is set up the tracing/logging/metrics exporters to directly speak with the sink. These days, it typically means specifying the URL and the DSN header.
- bilalq 2mo agoPerhaps there's a gap in my understanding. Can you clarify on this a bit more? I run a mix of serverless and non-serverless workloads. Gateway collectors are unavoidable because various SaaS platforms require you to be running publicly reachable endpoints to send telemetry to. In a runtime like Lambda, how would you avoid the need to run an edge collector? The only thing that comes to mind is to write to logs and then have a log stream processor that then writes to your gateway collector. Other than that, it seems unavoidable, no? Sure, in something like Fargate you could go app to sink. But even that has its own tradeoffs.
- clintonb 2mo ago(I’m not the person you replied to, but have experience here.) I follow the [gateway deployment pattern](https://opentelemetry.io/docs/collector/deploy/gateway/ https://opentelemetry.io/docs/collector/deploy/gateway/). Everything sends telemetry to our gateway, which exports to ClickHouse (formerly Datadog). We use Node.js, so all we need to do is run a script initializing Otel before running the app. We set this up following the docs a few years ago, and haven’t had to change it much since then.
- cute_boi 2mo agoI wish otel was never there. It is badly designed abstraction and due to otel the code gets very very messy and bad.
- ATMLOTTOBEER 2mo ago[flagged]
- EdSchouten 2mo agoWhat always puzzles me about OpenTelemetry is that tracing, metrics and logs are all designed independently. I wish there was a way I could just annotate my code base once, and let the ultimate decision to expose something as a metric/log/trace be dynamic at runtime. For example, if I look at a graph in monitoring dashboard and see something suspicious, I’d like to say: “The next time something like this occurs again, please save me a trace.” I should be able to just do that with a single mouse click. I remember them releasing the tracing spec/SDKs and saying “now let’s move on to metrics/logs.” That never sat right with me.
- thorian1828i03 2mo agoOne strategy do to do that is to trace everything by default and select what to sample later, e.g. https://grafana.com/docs/grafana-cloud/observe-and-act/adaptive-telemetry/adaptive-traces/ https://grafana.com/docs/grafana-cloud/observe-and-act/adapt...
- EdSchouten 2mo agoIf I understand that correctly, it means your app always creates traces, and Grafana Cloud is responsible for sampling/aggregating. That may be prohibitively expensive in terms of CPU/network load. What I’m suggesting is that your apps by default only send metrics to your monitoring system, but that the monitoring system can specifically ask to “upgrade” metrics to traces. Or to log entries. The same thing with metric cardinality: by default, only report metrics in a fully aggregated manner. But do tell the monitoring system how they can potentially be broken up if needed (i.e., which labels to add).
- thorian1828i03 2mo ago> The same thing with metric cardinality: by default, only report metrics in a fully aggregated manner. But do tell the monitoring system how they can potentially be broken up if needed (i.e., which labels to add). How does the monitoring system have any of the context to add labels? That would only exist in application memory. Grafana went the other way - your app exports all labels, and then you selectively aggregate on ingest: https://grafana.com/docs/grafana-cloud/observe-and-act/adaptive-telemetry/adaptive-metrics/ https://grafana.com/docs/grafana-cloud/observe-and-act/adapt... > That may be prohibitively expensive in terms of CPU/network load. In practice I've not experienced this even on quite high request rates. While it isn't free, exporting everything has been cheap enough that the real cost in dollars spent is basically marginal (it's _storing_ the data that's expensive)
- 0xbadcafebee 2mo agoIt is crazy to me how often people don't grok how to design software well. 1. The worst thing you can do is try to stuff too many things into one specification. So you want an API? That's great. What's that? You want a rigid set of types so that any tiny changes over time aren't compatible? You want to try to define every conceivable use case as a new call? You want to combine multiple elements from different domains into one flat set of functions? You don't have any hierarchy or inheritance? You don't support extensions? 2. The second-worst thing you can do is to force a whole lot of different people to go through a single standards body. So you want to support a thousand different 3rd party components. What's that? You want to require everyone get their adapter approved by one group? And there's only one supported adapter per 3rd party component? If you're trying to feed an entire city, it's logistically incredibly difficult to try to do it all yourself. If instead you just define where food can be dropped off or picked up, and ask volunteers to bring their own food there whenever they can/want, now you don't have a logistical nightmare on your hands anymore. The tech alternative? Add support for "plugins", make the plugin interface incredibly loose/backwards-compatible/layered, and invite people to publish their own plugins. If you under-engineer it, it actually works better.
- cyberax 2mo agoI disagree. I'm an observability geek, and OTel is... fine. It's missing a few things that I'd like, but I was able to implement them myself. I guess the major design issue is that the sampling decision is made at the _start_ of the segment. So I hacked up a few improvements: 1. Ability to mark segments as "boring", so they are dropped before the export. For things like healthchecks, empty "get the pending jobs" queries, etc. 2. Ability to downgrade errors for segments that are expected to return an error (e.g. HEAD on a non-existing object in S3 to check if there's a cached blob).
- masterj 2mo agoHN always grumbles about OTel, but I agree. It's fine, and important: https://jeremymorrell.dev/blog/opentelemetry-and-the-value-of-standards/ https://jeremymorrell.dev/blog/opentelemetry-and-the-value-o... I understand the author's perspective in the linked article, but none of that data shows a project in trouble? Some languages have more resources than others, but those all look like healthy open source projects
- steerpike 2mo agoOh my god. A Jeremy Morrell sighting in the wild. Every time I share your blog (and I share it a lot) I tell people: "This guy started a blog in 2024. Wrote three posts and all three of them would still make my top ten list of 'greatest posts on observability' today". 'A practitioner's guide to wide events' especially is still my number 1.
- masterj 2mo agoD'aww, thank you! I'm hoping to find time to write more this year
- jauntywundrkind 2mo agoI'd make a wager that things would go better smoother faster if folks tried more stuff, ventures forth more on their own. It's obviously not great that there's no semantic convention that's perfect and just works for everything, and yeah it takes a while. I feel like the real data I'd want is who else, how many people show up to say they've tried something. Is that happening? Whether specs are really good enough advance or not, to me, is often whether enough people have tried it to find out. The net of this is, otel is a very flexible system you can use and adapt in all kinds of ways and while the spec is important, using the toolkit to FAFO yourself, ahead of any beaten path, should really be encouraged. That's the message I'd want to see being radiated out about otel.
- rcleveng 2mo agoSounds a lot like K8s. It's not a framework you use, it's a framework to build a framework on top of. I wish the observability vendors would move to using it under the covers so it's easier to mix and match. I wish the otel support wasn't super buggy in most of the frameworks and backends.
- NegativeLatency 2mo agoBut then you wouldn’t be locked in!
- quadrifoliate 2mo ago[dead]
- osener 2mo agoI like the end result of OpenTelemetry tracing when using Axiom and the like, but the SDKs have been a nightmare. Too much emphasis on automatic instrumentation, Java-isms, everything is stateful and abstracted away. It can do distributed tracing of otherwise traditional long running microservices, but breaks down when your functions are distributed like in durable execution engines, Cloudflare Workflows, “functions” that span hours/days/weeks and steps that retry many times. I had to reverse engineer how SDKs work and how tracing UIs display data so I could make simpler functions that fit wider variety of runtimes and more freely parent spans, start spans and end them from different function instances. I think most of the API and terminology complexity is self inflicted. Would love to see a rebooted developer experience that is less Kubernates-brained.
- phrotoma 2mo agoI tried to emit metrics from a python app using otel once. Gave up and switched to prometheus. What a nightmare.
- jcmfernandes 2mo agoRan into the same issue and didn't find any willingness in the OTEL gods to close this gap.
- kalkin 2mo agoThe article assumes the issue with OTel is slow feature development, which isn't my experience at all. The issue I've had is that the SDKs have terrible performance overhead for instrumentation and are, as you say, highly resistant to integrating the output of better performing (or just preexisting) instrumentation. In Python and Ruby, at least, the CPU cost of all the mandatory abstraction is way too high.
- growse 2mo agoWhat I find confusing about this is that otel is two things. 1. A spec 2. A ref implementation Similar to other projects (e.g. python), if there's complaints about (2), that should trigger an ecosystem of alternative implementations that are guaranteed to be compatible because of (1). I suspect there's actually quite a few private, separate otel implementations. Maybe these just aren't being contributed as oss?
- gertburger 2mo agoI've found their django instrumentation to be kinda useless for larger apps. The only choices you get is full auto instrumentation, which breaks most non-trivial apps, or zero assistance/documentation. There is no in-between where I can inject the functionality required in a way that is compatible with the application.
- rm 2mo agoCould you please elaborate a bit on what is not working for you?
- gertburger 2mo ago(I haven't attempted to use opentelemetry-instrumentation-django in at least a year so my information might be dated and my memory is patchy :P) If I recall the primary issue was the forced loading of the django settings file by otel. I get that fully automated instrumentation should be turn-key and the current approach kinda works on basic applications. But most production django applications are monoliths and generally larger apps. They have non-trivial configuration processes which are often multi step and source settings from multiple places. Otel should not assume it can just randomly load a the django settings at an arbitrary time point in the startup process. In one of our apps the MIDDLEWARE setting specifically is dynamically generated and re-ordered based on enabled features. That application's startup process also has multiple stages and the initialisation of django occurs much later, after dependant config loaders etc have been initialised. What would allow us to integrate with opentelemetry-instrumentation-django much more easily is a set of smaller primitives that we can configure and call at the appropriate time. opentelemetry-instrumentation-django has (had?) a lot of logic hidden inside a large "inject" function which could not easily be extracted into the constituent parts and applied in a compatible manner. https://github.com/open-telemetry/opentelemetry-python-contrib/issues/2301 https://github.com/open-telemetry/opentelemetry-python-contr...
- rm 2mo agoThanks for the write up, appreciated. A couple of things: - users are not forced to use auto-instrumentation. People can import the Middleware and use it as they see fit. I see that the instrumentor is configuring the middleware using some private attributes, I guess that can be extracted into a public function so it would be easier to do so - speaking of the middleware, the chances that it'll become a public symbol are scarce as are the chances that the interfaces will change. So if one has some testing before going to production it should be fine
- tete 2mo agoOpenTelemtry is the perfect example of an overengineered mess. While I usually think that at least having some standard that people agree on I think OpenTelemtry should be dropped. A lot of the less popular alternatives (just going with Prometheus, Victoriametrics, etc) are de-facto competing smaller standards and a lot better both in terms of less added complexity and the results you get. I think OpenTelemetry turned metrics into a farce. In many situations even self-rolled telemetry works better even with the added stuff. The annoying thing is that OpenTelemtry is that big standard now one kind of has to to add compatibility. So please, if you write software, make sure you don't lock yourself into OTel.
- nlitened 2mo agoI agree overall, however: > A lot of the less popular alternatives (just going with Prometheus, Victoriametrics, etc) are de-facto competing smaller standards By all metrics (hah), Prometheus is the more popular solution and is the de-facto standard, as far as I know.
- dengolius 2mo agoWhat about InfluxDB? Speaking of standards, it’s worth noting that this system was created early on, when microservices were emerging as a concept. Another reason this “de facto standard” emerged is that there were almost no alternatives. This is what we need to understand about standards and who promotes them. But technology doesn’t stand still, and I wouldn’t call Prometheus the standard right now, because OpenTelemetry is already starting to be referred to as the standard in the field of observability. Some people like it, some don’t, but Zabbix and Nagios are still going strong; for some, they remain the standard for monitoring for a variety of personal reasons.
- time4tea 2mo agoIts a shame that the various implementations are pretty horrible. Global state, static methods etc etc. If you get rid of that, and just pass dependencies around, create some appropriate local abstraction around them.. the tooling, be it datadog or honeycomb does a great job making it useful. Can't really say the same for grafana, but ymmv - depending on budget
- jgalt212 2mo agoPremature instrumentation is the root of all evil. And the source of a significant part of AWS revenue. It should not cost more to monitor an app then run it.
- ishan_vats 2mo ago[flagged]
- Havoc 2mo agoI find the entire observability space to quite a poor experience, at least in the self-hosted space. Tried both grafana route and signoz and neither seems particularly pleasant
- N_Lens 2mo agoTry datalust/seq
- nunez 2mo agoWhat about the experience did you find lacking?
- Havoc 2mo agoThat would be an essay, but in short: Grafana is fine to get to a selfhosted basic install. But once you try to actually connect logs, metrics, traces in selfhosted context and perhaps sprinkle some otel in...that sht gets out of hand very fast. It's modular in a way that seems like a win but once you start connecting stuff it starts adding complexity not ease. Signoz...still pretty early in exploring this and so far it's acceptable, but it too relies on a mix of query languages incl the competitors promql so some panels support it other seem to not to? The entire thing just seems bewildering to me. Not gonna say "why is this so hard" because I genuinely thing smart people are genuinely trying here...but there result just isn't great.
- ruhani_grover 2mo ago[dead]
- dengolius 2mo ago> Grafana is fine to get to a selfhosted basic install. But once you try to actually connect logs, metrics, traces in selfhosted context and perhaps sprinkle some otel in...that sht gets out of hand very fast. It's modular in a way that seems like a win but once you start connecting stuff it starts adding complexity not easy. Maybe because some open-source software has limitations? That seems reasonable to me.
- greatgib 2mo agoI have always been turned off to attempt to use OTel by the feeling that it is a little bit too over-engineered a that it might be very bad in term of performance/wasted network traffic when you see the data structure that it is using.
- deleted 2mo ago[deleted]
- huksley 2mo agoOTel is very complicated while yeah for example datadog is just dropin. And Graylog support for OTel makes it a second class citizen in the logs (all attributes are prepended with otel_attributes_ which makes searching difficult). Using is hard, vendors are hostile, it seems like no-one want it to be a first class citizen...
- Kinrany 2mo agoIt feels like OTel tried standardizing before the correct design was anywhere close to being settled. It's only time to standardize once there's consensus on all the important points, and what's left is minor details that don't matter for anything other than compatibility.
- ninkendo 2mo agoSpeaking only from my experience using their rust crates, they have undergone more “code feng shui” than any of our other dependencies. They’re still 0.x and every point release seems to re-imagine things enough to break everything and require substantial rewriting. They don’t even bother describing the motivation for changes, just, you can’t use this type any more, it’s private now. You can’t configure metadata here any more, you have to do it there now. It’s been the most painful dependency of ours by far.
- dijit 2mo agoI know sadly very little about otel, it feels “heavy” in a way I am not used to, I am used to simple systems - configured and composed in a way that makes a larger system. 20 years ago, we were doing (what I think) OTel is doing: with “hit IDs” (half way between a session and a request) that were consistently applied when logging the cause a request being fired; along centralised logging and really good timekeeping. Essentially a unique identifier as a tag that followed the request as it passed through the system. This was enough to debug basically any problem. We could even measure the distance between requests of the same “hit” and the total wall-time before it managed to return through the load balancer, so we could track our p99 easily. Though truthfully we didn't make pretty graphs. I sometimes wonder what OTel gives me more than this, but I work in games now and lots of these things that work well in webdev do not apply at all to our problems.
- masterj 2mo agoYou are essentially describing a proto-tracing system. At the risk of self-promoting twice in one comments section, I have a post walking through going from what you describe above to OTel-compatible tracing: https://jeremymorrell.dev/blog/minimal-js-tracing/ https://jeremymorrell.dev/blog/minimal-js-tracing/ You are right that what you were doing is very similar! However standardization helps a lot here.
- czhu12 2mo agoIt really never grokked with me why there isn't just "open source Datadog" that can be installed and used. End to end, stateful, that we can just self host. Our team tried to set up open telemetry to replace Datadog and got totally crushed in complexity. The model of having Open Telemetry just be for standardizing & exporting to other backends, needing glue for each part of the setup was nuts.
- reactordev 2mo agoSignoz? But yes it seemed like OTel was more interested in being a spec than a tool.
- frez1 2mo agoisn't this exactly what the LGTM stack is?
- losingthefight 2mo agoI run OSS Grafana with Loki, Prometheus, and Tempo. I use an Alloy sidecar taking in OTEL and scraping logs.feom my Go services and selfhost the stack. Once you need to scale it gets a bit more complicated but it's all still OSS. The biggest challenge I have is that each data source needs it's own query language, which DD and the like don't. That's why at my day job they went with DD despite the costs. Still OTEL but the querying is the same. We are also looking at Dash0 but for all of my personal and consulting jobs, OSS LGTM/P works good for me.
- pphysch 2mo agoThere is, it's called VictoriaMetrics/Logs/Traces. https://victoriametrics.com/ https://victoriametrics.com/
- rtpg 2mo agoWe use victoriametrics, but I believe that's just the collector side of it cuz we also query into it with grafana. Datadog isn't just a collector, but the whole querying UI as well, right?
- dwoldrich 2mo agoI think the industry would benefit from some general evangelism for observability. Being able to do distributed tracing was both a "well, duh" and mindblown experience when I first learned about it a decade ago. It made supporting software so much better. OTel is a fine system for learning observability; it does an okay job of exposing capabilities given how diverse the vendor ecosystem is.
- tablloyd 2mo ago> However on the collector side you end up having to do the OpenTelemetry Collector Builder to make your own collector (or just kinda ride the wave and hope it works out). While cool that this exists, it's a lot of scope to ask a team to take on. This is just plain wrong, binaries of the collector are shipped which are available to use straight away. You can use the builder if you want to create your own version with a selected set of components but it is no way a hard requirement.
- psadri 2mo agoI looked at the OTel schema generated for recording a single numerical metric. It was like 12 or 13 meta fields in addition to the actual metric fields like timestamp, metric and value. OTel looks like something designed by a committee of committees, funded by someone who is in the business of selling cloud storage/data warehousing services.
- mnming 2mo agoI also think OTel SDKs are a tad bit too prescriptive, but at the same time I can't envision what a better version would look like. The core of logs and spans are just wide events with some inter-connections, those SDKs and OTel docs make them less obvious. (I maintain o11ylite https://github.com/o11ylite/o11ylite https://github.com/o11ylite/o11ylite)
- awill88 2mo agootel is amazing and when thoughtfully instrumented, turns out, you can opt out of auto instrumentation btw, it provides a standard that is useful, well maintained, portable to enterprise or self hosted. It’s modular, but the author is appraising the Ruby shortcomings as a problem while also saying they unfortunately don’t have time to contribute because, you know, they can’t “join the calls” lol We just got a CTO who loves Ruby and guess what I’m about to do: use AI to fill in the Ruby gaps and open a PR and work a weekend or two and see if they like it and then you won’t write any more articles disparaging a project that I personally love. It saves our company AT LEAST 10k a month vs having datadog / splunk / enterprise-y bullcrap. You seem so educated, why not roll up your sleeves instead of patronizing the hard working people that make the project work with your “if it were me” just go ahead and say it in their forums. And, if you work for your paycheck, you’re using an agent. A mature project like Otel? Shit. That’s easy-peasy to feed into an agent, so what’s what is actually the problem? Take the time to learn and help them out if it bothers you so much you want to share it with the world! Isn’t every single “problem” found in every large and successful open sourced framework? Throwing out a baity framing like it’s some kind of project going wrong and then kind of just ending the article without making any sort of judgement on where this all leads, proceeding to post on hackernews.. bait! It’s a cloud native project that is not owned by any company. That’s so rare and worth an article to celebrate open source! What a privilege to stand on the shoulders of giants! > So OpenTelemetry currently is attempting to support a dizzying number of languages and frameworks. “dizzying” — so I’m lost, did the author remmeber the scope of the project before they started making judgements about it? And calling the attention of hackernews here: what’s the alternative? Oh that’s right, there isn’t one. Because this is a wag my finger article for attention and aggregating the author on a developer channel to boost their presence. Lame. (Thumbs down)
- suralind 2mo agoI don't understand the sentiment. OTEL is better than anything I've ever used before. Do I like every part? No. My personal no-no is the automatic instrumentation which I always bypass and just DI it myself, I don't like the "global" by default approach in Golang and I had to fight team mates who were all for using it. That said, no other observability library that I've ever used was so good overall. The perf is meh, but tbh if you look at the kind of code we, regular developers write for work, it's probably still vastly better.