30 ms·
OpenTelemetry Tracing in < 200 lines of code
- thewisenerd 2y agothe "true spec is the data" is very powerful. for example, we translate our loosely OTEL-based telemetry into a format which was consumable by any otel collector. shim a few fields, et voila! can be read by Jaeger-UI. free trace tree visualization.
- alisonatwork 2y agoI agree. The hardest work on OpenTelemetry (and OpenCensus/OpenTracing before it) was looking at all the different vendors and trying to come up with a common set of semantic conventions[0]. If your team is new to metrics or tracing - or even just structured logging - it's worth to start adding fields in the general structure of the otel semantic conventions, because then whichever third party service you eventually decide to push to, it won't take much of a shim to adapt your data to get there. And if you just stick with JSON logs pushed to ELK (or whatever) you at least build up a useful set of fields to query on. [0] https://opentelemetry.io/docs/concepts/semantic-conventions/ https://opentelemetry.io/docs/concepts/semantic-conventions/
- Ciantic 2y agoThe utility of tracing is great, I've been using Azure Application Insights with NodeJS (and of course in .NET). This is relatively simple because it monkey patches itself everywhere if you go through the "classic" SDK route. Then adding your own data to logs is just a few simple functions trackTrace, trackException, trackEvent, etc. However, if you want to figure out how it works you might be scared, it is not lightweight. I just spent a few days digging through the Azure Application Insights NodeJS code base which integrates with OpenTelemetry packages. It's an utter mess, a huge amount of abstractions. Adding it to the project brought 100 MB and around 40 extra packages.
- chucklenorris 2y agoYes, this is exactly my impression too.. the code for opentelemetry-js is over engineered and adds a scary amount of dependency code. There are quite a bunch of libraries which I'm not sure what they do and in which scenarios I might need them. The documentation is not very helpful either. I look forward to someone implementing a opentelemetry-nano package with only the minimum stuff needed and allow me to choose extra support for my dependencies or an easy way of adding my own wrappers.
- pimeys 2y agoAlso badly documented. If you try to implement something non-standard with it, good luck. I once needed to write code where trace started in node an continued inside a node api native library. Getting these two traces to connect must be one of the most frustrating things I've worked on. At least on the Rust side you have types to help you out, but it is still quite complex and the crates have bugs open for years, impossible to solve with the current architecture.
- wordofx 2y agoWhat’s your plans for applications insights sunsetting?
- MuffinFlavored 2y agohttps://azure.microsoft.com/en-us/updates/we-re-retiring-classic-application-insights-on-29-february-2024/ https://azure.microsoft.com/en-us/updates/we-re-retiring-cla... Do you have a link for what you are speaking of?
- wordofx 2y agoThere’s no public announcement yet but from what reps say to customers and what people working on azure say is app insights is more or less being wound down in favor of building out open source solutions because it’s more favourable and less maintaince / dev than building out their own solution. Think more OTEL/Grafana. Basically word on the inside is MS doesn’t want to pay to build out app insights.
- h1fra 2y agoI have one biff against otel, it's not possible to stream a trace. You have to build a big object in memory that is not suitable for a long-running process. And because of that it's not possible to start it somewhere and finish it somewhere else
- thephyber 2y agoI'm under the impression this is wrong. A trace is made from one or more spans. So long as the context is propagated[1] from one service to another, any number of spans referring to the same trace can be generated anywhere at any time. The trace+spans don't have to be created in the order of a stack. [1] https://opentelemetry.io/docs/concepts/context-propagation/ https://opentelemetry.io/docs/concepts/context-propagation/
- phillipcarter 2y agoThis is correct. Traces are made up of spans that can be created within the same process, different processes, different machines, etc. and all emitted asynchronously.
- eterm 2y agoI don't think that's a limitation of the specification. You can create spans and emit events for those spans without holding objects in memory.
- malkia 2y agoThere is no concept of streaming trace. You are propagating a random ID, and expect other NODES (yours, or outside of your control) to re-emit these, and if need to create new ones. The agreement is that these nodes would either emit all these events to (eventually) a common place (push), or something is going to gather them (pull). I think at Google, some of these used to be still on the machines, and the tools would pull them directly, and back then it was possible to mark certain for preservance (that was long time ago - 2014, so I'm sure things have changed). Also I was on the over-excited edge, because I was not aware what was this, but I was on call (a small team in ads), and had to page up to the Bigtable/Megastore or was it Spanner team, and they simply asked me to bump some tracing bits up for like 30 seconds, then something magically showed up - and I was - wtf! I think it was then it clicked with me how useful this is, .... but also how much wasteful (in terms of resources) it could be if you don't end up looking there.
- krashidov 2y agoIt's so easy in node. I miss node. Setting this up in the mess that is Python/Gunicorn/asgi/wsgi/celery/Django has not been as easy
- tempest_ 2y agoSentry has been nice. I do not know what unholy monkey patching they do with that sentry_sdk.init call and I try not to think about it but for web apps it is fire and forget.
- krashidov 2y agohmm do you use Sentry for logging? We use sentry as well but only for errors. Also the trace ids don't match with the logging traces so I have to fix that too
- optiomal_isgood 2y ago> I do not know what unholy monkey patching they do A year ago I had to patch its package for an internal use case. Their codebase is fairly well-written I thought (at least the JS SDK). e.g. this is where they add the `fetch` breadcrumb https://github.com/getsentry/sentry-javascript/blob/develop/packages/browser/src/integrations/breadcrumbs.ts#L312 https://github.com/getsentry/sentry-javascript/blob/develop/... e.g. where the actual monkey patching happens https://github.com/getsentry/sentry-javascript/blob/e1783a653fafda6df9eb55e1eaf61e113f5df3db/packages/utils/src/instrument/fetch.ts#L78 https://github.com/getsentry/sentry-javascript/blob/e1783a65...
- tempest_ 2y agoMy experience in principally in python and the number of integrations that are auto used on common dependent libs and frameworks is long https://docs.sentry.io/platforms/python/integrations/ https://docs.sentry.io/platforms/python/integrations/
- viraptor 2y agoYou can do almost 1:1 the same thing (compared to the post) in Python and use it as a wsgi/whatever middleware. It's really not any different. The callback changes to a resource manager, but that's about it.
- caseyw 2y agoOpenTelemetry (OTel) and the OpenTelemetry Protocol (OTLP) are immensely powerful tools. The ability to emit telemetry data from any source, coupled with a receiver that can sample, filter, pipe, and potentially reshape the data to suit any need, is a game changer. This flexibility revolutionizes how we approach observability and monitoring across diverse systems.
- vrosas 2y agoWhile the libraries and the documentation for otel are bloated messes, I maintain that any platform that isn’t using some sort of tracing system is practically negligent in their engineering duty. If you’re still out there querying logs with some giant sql statements you’re missing out. The pure wonder of being able to click on an http request and seeing every service it touched, every application log outputted, and every database query it ran, and all the timings of each of those is magical.
- KronisLV 2y ago> I maintain that any platform that isn’t using some sort of tracing system is practically negligent in their engineering duty. For some, it's difficult because many of the self-hostable out there are rather complex and have high requirements, like https://github.com/getsentry/self-hosted/blob/master/docker-compose.yml https://github.com/getsentry/self-hosted/blob/master/docker-... Personally I found Apache Skywalking to be something that you can setup without too many issues https://skywalking.apache.org/ https://skywalking.apache.org/ but it's not exactly ideal either. I wonder what other good options are out there, something that you can have up and running on a 5$ VPS within an hour or two, to not cause friction. Where's the OpenTelemetry equivalent of launching an (opinionated) Docker Compose stack that has everything you need on the server side, running against SQLite, MariaDB, PostgreSQL, ClickHouse, ElasticSearch or another data store? Of course, when SaaS is an option, many will just go for that.
- vanschelven 2y agoI wrote https://bugsink.com/ https://bugsink.com/ especially because of the complexity and high requirements of self-hosted sentry, but the focus is on Error Tracking rather than being a full APM-like solution
- dipakparmar 2y agoThanks for sharing your insights! I agree that having a good tracing system is essential for modern engineering practices. Skywalking is indeed an interesting choice with a decent feature set, especially given its open-source nature and support for various integrations. I would love to hear more about your experience with it. Have you explored any other tools, like Grafana OSS? Grafana has a robust stack that many find easy to set up using Docker or Docker Compose. It offers the flexibility to run in single-container mode or to scale to high-availability clusters, which is a big plus. For me, having control over the platform is crucial, and I appreciate solutions that allow a smooth transition to their SaaS counterparts without the hassle of data migration. I'm curious to know what specific features of Skywalking stand out to you and what you're hoping to achieve with a tracing system
- sandelz 2y agoWhile otel is really nice and easy to integrate (at least on .net and node) into software, the collector/UI side seems to be overly complex. I have used application insights on Azure on my day job but I was wondering is there a simple self hosted collector/UI to use?
- WuxiFingerHold 2y agoI've done something similar for a small service running in Docker compose. Maybe I didn't think through it but the complexity and huge dependencies introduced by the official libs immediately turned me off. DIY was very easy and simple.