6 ms·
GitHub CI/CD observability with OpenTelemetry step by step guide
- candiddevmike 1y agoHow does SigNoz compare to the other "all-in-one" OTel platforms? What part of the open-core bit is behind a paywall?
- makeavish 1y agoOnly SAML, Multiple ingestion keys and Premium Support is under paywall. SSO is not under paywall. Check pricing page for detailed comparison: https://signoz.io/pricing/ https://signoz.io/pricing/
- drzaiusx11 1y agowe run a prom+otel+xray stack at scale with grafana as the interface. i honestly don’t miss datadog at all at this point. trivial to add new alarms, dashboards etc
- reactordev 1y agoAs someone who has some experience in observability at scale, the issue with SigNoz, Prom, etc is that they can only operate on the data that is exposed by the underlying infrastructure where the IaaS has all the information to provide a better experience. Hence CloudWatch. That said, if you own your infrastructure, I’d build out a signoz cluster in a heartbeat. Otel is awesome but once you set down a path for your org, it’s going to be extremely painful to switch. Choose otel if you’re a hybrid cloud or you have on premises stuff. If you’re on AWS, CloudWatch is a better option simply because they have the data. Dead simple tracing.
- 6r17 1y agoI did have some bad experiences with OTEL and have lot of freedom on deployment ; I never read of Signoz will definitely check it out ; SigNoz is working with OTEL I suppose ? I wonder if there are any other adapters for trace injest instead of OTEL ?
- darkstar_16 1y agoJaeger collector perhaps but then you'd have to use the Jaeger UI. Signoz has a much nicer UI that feels more integrated but last I checked had annoying bugs in the UI like not keeping the time selection when I navigated between screens.
- 6r17 1y agoDefinitely should look up the tech more ; i lazily commented as Signoz clearly state it ingest most than 50 different sources ;
- elza_1111 1y agoyep, SigNoz is OpenTelemetry native. You can instrument your application with OpenTelemetry and send telemetry data direclty to signoz.
- bbkane 1y agoThere are a few: I've played with https://uptrace.dev https://uptrace.dev and https://openobserve.ai/ https://openobserve.ai/ . OpenObserve is a single binary, so easy to set up
- mdaniel 1y agobe cognizant of their licenses (AGPLv3), it matters in some shops https://github.com/uptrace/uptrace/blob/v1.7.6/LICENSE https://github.com/uptrace/uptrace/blob/v1.7.6/LICENSE https://github.com/openobserve/openobserve/blob/v0.14.7/LICENSE https://github.com/openobserve/openobserve/blob/v0.14.7/LICE...
- FunnyLookinHat 1y agoI think you're looking at OTel from a strictly infrastructure perspective - which Cloudwatch does effectively solve without any added effort. But OTel really begins to shine when you instrument your backends. Some languages (Node.js) have a whole slew of auto-instrumentation, giving you rich traces with spans detailing each step of the http request, every SQL query, and even usage of AWS services. Making those traces even more valuable is that they're linked across services. We've frequently seen a slowdown or error at the top of our stack, and the teams are able to immediately pinpoint the problem as a downstream service. Not only that, they can see the specific issue in the downstream service almost immediately! Once you get to that level of detail, having your infrastructure metrics pulled into your Otel provider does start to make some sense. If you observe a slowdown in a service, being able to see that the DB CPU is pegged at the same time is meaningful, etc. [Edit - Typo!]
- makeavish 1y agoAgree with you on this. OTel agents allows exporting all host/k8s metrics correlated with your logs and traces. Though exporting AWS service specific metrics with OTel is not easy. To solve this SigNoz has 1-Click AWS Integrations: https://signoz.io/blog/native-aws-integrations-with-autodiscovery/ https://signoz.io/blog/native-aws-integrations-with-autodisc... Also SigNoz has native correlation between different signals out of the box. PS: I am SigNoz Maintainer
- elza_1111 1y agoFYI for anyone reading, OTel does have great auto-instrumentation for Python, Java and .NET also
- reactordev 1y agoNot confusing anything. Yes you can meter your own applications, generate your own metrics, but most organizations start their observability journey with the hardware and latency metrics. Otel provides a means to sugar any metric with labels and attributes which is great (until you have high cardinality) but there are still things that are at the infrastructure level that only CloudWatch knows of (on AWS). If you’re running K8s on your own hardware - Otel would be my first choice.
- elza_1111 1y agoThere are integrations that let you monitor your AWS resources also on SigNoz. That said, I personally think CloudWatch is painful in so many other ways as well, Check this out, https://signoz.io/blog/6-silent-traps-inside-cloudWatch-that-can-hurt-your-observability/ https://signoz.io/blog/6-silent-traps-inside-cloudWatch-that...
- mdaniel 1y agoA child comment mentioned k8s but I also have been chomping at the bit to try out the eBPF hooks in https://github.com/pixie-io/pixie https://github.com/pixie-io/pixie (or even https://github.com/coroot/coroot https://github.com/coroot/coroot or https://github.com/parca-dev/parca https://github.com/parca-dev/parca ) all of which are Apache 2 licensed The demo for https://github.com/draios/sysdig https://github.com/draios/sysdig was also just amazing, but I don't have any idea what the storage requirements would be for leaving it running
- bravesoul2 1y agoThat's a genius idea. So obvious in retrospect.
- hrpnk 1y agoHas anyone seen OTel being used well for long-running batch/async processes? Wonder how the suggestions stack up to monolith builds for Apps that take about an hour.
- zdc1 1y agoI've tried and failed at tracing transactions that span multiple queues (with different backends). At the end I just published some custom metrics for the transaction's success count / failure count / duration and moved on my with life.
- madduci 1y agoI use Otel running in a GKE cluster and tracking Jenkins jobs, whose spans/traces can track long time running jobs pretty well
- makeavish 1y agoYou can use SpanLinks to analyse your async processes. This guide might be helpful introduction: https://dev.to/clericcoder/mastering-trace-analysis-with-span-links-using-opentelemetry-and-signoz-a-practical-guidepart-2-1amc https://dev.to/clericcoder/mastering-trace-analysis-with-spa... Also SigNoz supports rendering practically unlimited number of spans in trace detail UI and allows filtering them as well which has been really useful in analyzing batch processes: https://signoz.io/blog/traces-without-limits/ https://signoz.io/blog/traces-without-limits/ You can further run aggregation on spans to monitor failures and latency. PS: I am SigNoz maintainer
- ai-christianson 1y agoIs this better than Honeycomb?
- mdaniel 1y ago"Better" is always "for what metric" but if nothing else having the source code to the stack is always "better" IMHO even if one doesn't choose to self-host, and that goes double for SigNoz choosing a permissive license, so one doesn't have to get lawyers involved to run it --- While digging into Honeycomb's open source story, I did find these two awesome toys, one relevant to the otel discussion and one just neato https://github.com/honeycombio/refinery https://github.com/honeycombio/refinery (Apache 2) -- Refinery is a tail-based sampling proxy and operates at the level of an entire trace. Refinery examines whole traces and intelligently applies sampling decisions to each trace. These decisions determine whether to keep or drop the trace data in the sampled data forwarded to Honeycomb. https://github.com/honeycombio/gritql https://github.com/honeycombio/gritql (MIT) -- GritQL is a declarative query language for searching and modifying source code
- sali0 1y agonoob question, i'm currently adding telemetry to my backend. I was at first implementing otel throughout my api, but ran into some minor headaches and a lot of boilerplate. I shopped a bit around and saw that Sentry has a lot of nice integrations everywhere, and seems to have all the same features (metrics, traces, error reporting). I'm considering just using Sentry for both backend and frontend and other pieces as well. Curious if anyone has thoughts on this. Assuming Sentry can fulfill our requirements, the only thing taht really concerns me is vendor-lockin. But I'm wondering other people's thoughts
- whatevermom 1y agoSentry isn’t really a full on observability platform. It’s for error reporting only (that is annotated with traces and logs). It turns out that for most projects, this is sufficient. Can’t comment on the vendor lock-in part.
- srikanthccv 1y ago>I was at first implementing otel throughout my api, but ran into some minor headaches and a lot of boilerplate OTeL also has numerous integrations https://opentelemetry.io/ecosystem/registry/ https://opentelemetry.io/ecosystem/registry/. In contrast, Sentry lacks traditional metrics and other capabilities that OTeL offers. IIRC, Sentry experimented with "DDM" (Delightful Developer Metrics), but this feature was deprecated and removed while still in alpha/beta. Sentry excels at error tracking and provides excellent browser integration. This might be sufficient for your needs, but if you're looking for the comprehensive observability features that OpenTelemetry provides, you'd likely need a full observability platform.
- dboreham 1y agoYou can run your own sentry server (or at least last time I worked with it you could). But as others have noted sentry is not going to provide the same functionality as OTel.
- mdaniel 1y agoThe word "can" is doing a lot of work in your comment, based on the now horrific number of moving parts[1] and I think David has even said the self-hosting story isn't a priority for them. Also, don't overlook the license, if your shop is sensitive to non-FOSS licensing terms 1: https://github.com/getsentry/self-hosted/blob/25.5.1/docker-compose.yml https://github.com/getsentry/self-hosted/blob/25.5.1/docker-...
- totetsu 1y agoI spent some time working on this. First I tried to make a GitHub action that was triggered on completion of your other actions and passed along the context of the triggering action in the environment, then used the GitHub api to call out extra details of the steps and tasks etc, and the logs and make that all into a process trace and send it via an otel connection to like jaeger or grafana, to get flamchart views of performance of steps. I thought maybe it would be better to do this directly from the runner hosts by watching log files, but the api has more detailed information.
- 127dot1 1y agoThat's a poor title: the article is not about CI/CD, it is particularly about GitHub CI/CD and thus is useless for the most CI/CD cases.
- dang 1y agoOk, we've added Github to the title above.
- ankit01-oss 1y agothanks dang
- remram 1y agoI have thought about that before, but I was blocked by the really poor file support for OTel. I couldn't find an easy way to dump a file from the collector running in my CI job and load it on my laptop for analysis, which is the way I would like to go. Maybe this has changed?
- sweetgiorni 1y agohttps://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/fileexporter https://github.com/open-telemetry/opentelemetry-collector-co...
- remram 1y agoAnd the receiver: https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/receiver/otlpjsonfilereceiver https://github.com/open-telemetry/opentelemetry-collector-co... I'll have to try this! edit: actually Jaeger can just read those files directly, so no need to run a collector with the receiver. This is great!