4 ms·
I think it would be great to show the performance impact of these SDKs because it is one of the really important aspects of monitoring (being non-intrusive).
by StreamBright 5y ago
I think it would be great to show the performance impact of these SDKs because it is one of the really important aspects of monitoring (being non-intrusive).
- crandycodes 5y agoAs someone who's needed to maintain complex, high-performance database drivers that needed to work across a bunch of different platforms, I've been following them and their predecessors of OpenTracing/OpenCensus. The problem that's always been interesting to me as a library maintainer is consistency across platforms and well maintained multi-platform libraries. I hadn't really found an acceptable solution that would work across Java, Node.js, browser, and so on. We'd invented our own formats and then we owned all the integration problems with various monitoring tools. I left the team before we started to adopt, but they've started doing it and it looks like it's help with reducing integration burden. I also think using someone else's opinionated library can help avoid bikeshedding on concepts not related to your core value.
- jeffbee 5y agoI've noticed that orgs where I've worked vary between being totally insensitive to observability cost to being real hardasses about it. But I think most smaller shops are falling into the former category. I've even heard in meetings crazy shit like "It's very low overhead, only about 5%" which would get you laughed out of the office at, say, Google. Unfortunately (to me) the focus on ease-of-use has meant that OpenTelemetry concepts are structured in such a way to preclude even the possibility of a very efficient implementation, which means that there will be a schism between people who are happy in the otel ecosystem and people who can't use it on cost grounds, who probably will splinter into distinct home-grown solutions.
- drewbug01 5y ago> Unfortunately (to me) the focus on ease-of-use has meant that OpenTelemetry concepts are structured in such a way to preclude even the possibility of a very efficient implementation Curious what you mean about the design of OpenTelemetry precluding efficient implementation?
- pm90 5y agoYou're correct that 5% increase in resource usage is probably not noticeable for most orgs. Its important to know what audience you're building for. I believe the audience for otel consists largely of companies that don't look anything like Google. So its fine to sacrifice that last bit of perf gain if it means the code is easier to use and maintain. FWIW, it also appears that companies like Google would fork or reimplement such systems anyway.
- legulere 5y agoIt’s also about where your costs are. Google has much less revenue per compute-time/requests/whatever than most other companies. If you target the business to business field your computing costs are usually negligible while your developers cost a lot of money. Throwing more hardware at the problem is often the most economical solution.
- ungzd 5y ago> Unfortunately (to me) the focus on ease-of-use has meant that OpenTelemetry concepts are structured in such a way to preclude even the possibility of a very efficient implementation Looks like Opentelemetry (at least its precedessor, OpenCensus) is originated from Google. From OpenCensus website (https://opencensus.io/ https://opencensus.io/): > OpenCensus and OpenTracing have merged to form OpenTelemetry > OpenCensus originates from Google, where a set of libraries called Census are used to automatically capture traces and metrics from services. Original internal Google tracing system was probably designed for scale. And opentelemetry's design is probably based on that internal system. So, maybe poor performance is just an implementations issue.
- jeffbee 5y agoOpenCensus bears no resemblance whatsoever to Google Census, except for the name. Census at Google is wired tight as hell. Any time it showed up on the first page of fleet-wide profiles it would get hammered back down. At the same time it also has more capabilities than its open source successors. Unfortunately, nothing has ever been published about it.
- austinlparker 5y agoI want to point out that OpenTelemetry is explicitly designed to decouple the API from the SDK, allowing for other people to reimplement parts of the SDK as needed while maintaining compatibility with not only other SDK components, but also the overall ecosystem. This was one of the major changes we introduced as part of the OpenTracing/OpenCensus merger.
- malkia 5y agoLook into envoy/istio - e.g. these introduce side-processes (sidecars) where your process talks to, and these create the traces for you at some perf cost. There are proxies for some of the existing services - like mysql - https://www.envoyproxy.io/docs/envoy/latest/configuration/listeners/network_filters/mysql_proxy_filter https://www.envoyproxy.io/docs/envoy/latest/configuration/li... I haven't used it myself, I've used census, and now looking into OpenTelemetry (though from the least finished version - C++). Had mixed success with it in the past, but trying again. Also not looking at all into side-cars, etc. - We compile all our internal tools, so adding this inside is where I'm getting into. I've had several times (while at Google), being asked by an SRE that I would call on issues, and they would request to bump the sample tracing from minimal defaults (was it 1 in 100,000 or million - forget) sometimes to 1:1 - for say 30 seconds. This way they'll receive on their end (in their systems that we use) flags to sample too, and at the end get full logs. Usually the whole trace is visible in few minutes. There were few UI's (nothign like zipkin/jaeged/others outside) - some of them with very "imgui"-like hacky (in good sense) view - like programmer art all over (which I loved - it was much more condensed than standard zipkin/jaeger). You could've marked something as important, and it'll retain for longer period - otherwise - poof - soon gone. Also it would collect info sometimes directly from the machine it was in (rather than wait to populate). Obviously, I don't know the details - I was just an user, or more like - allowing (when oncall) trace sampling to be bumped by the SRE - so they would get more info. It's what hooked me actually, because how else would one get everything end-to-end. Surprisingly it's also useful for single apps, where you have threads (or concurrency tasks, like with TBB/ConCRT) doing nested parallel_for's or spawning jobs, and you want to get idea what's going on. The only tricky bit is how to get your "context" propagated from one thread to another (also not readily done). It's one thing that the "golan" got right with their context for example. So it's really awesome, but probably really hard to get right the first few times.
- Thaxll 5y agoenvoy/istio do not replace telemetry in application because they only see what's pass through them and don't know anything else. You're missing a lot if you only instrument through a proxy.
- Philip-J-Fry 5y agoWell, the second you do any database call or other service call you've already spent 100x longer doing that than you have recording some timings. These clients will usually buffer the stats in memory and push them out asynchronously. Performance is definitely affected but I'm pretty sure it's negligible for most cases. Best practice would be to reduce tracing ratio in production too. So most requests are literally just a timing.
- sethammons 5y agoit is a solid call out. One of our teams had to rip out and completely re-think our integration with OpenTracing because of allocations in the client at the time. I believe that's been fixed.