10 ms·
Saving Three Months of Latency with a Single OpenTelemetry Trace
- darig 2y ago[dead]
- ochronus 2y ago[flagged]
- deleted 2y ago[deleted]
- simonbarker87 2y agoI wish posts like this would explore the relative savings rather than the absolute. On its own I don’t feel like that saving is really telling me much, taken to the extreme you could just not run the service at all and save all the time - a tongue in cheek example but in context is this saving a big deal or is it just engineering looking for small efficiencies to justify their time?
- tnolet 2y agoHey, I work at Checkly and asked my coworker (who wrote the post) to give some more background on this. I can assure you, we're busy and this was not done for some vanity price!
- janOsch 2y agoI'm the author of the post. You raise a good point about relative savings. Based on last week's data, our change reduced the task time by 40ms from an average of 3440ms, and this task runs 11 million times daily. This translates to a saving of about 1% on compute.
- simonbarker87 2y agoThanks for the follow up, sounds like a decent saving and investment of time then.
- janOsch 2y agoFun fact: it probably took more time to write up and refine the blog post than it did to hunt down that sneaky 40ms savings.
- simonbarker87 2y agoTrue but the value of the hunt and fix may really come from this blog post long term. Content marketing and all that
- hiatus 2y ago> This translates to a saving of about 1% on compute. Does this translate to any tangible savings? I'm not sure what the checkly backend looks like but if tasks are running on a cluster of hosts vs invoked per-task it seems hard to realize savings. Even per-task, 40 ms can only be realized on a service like Lambda—ECS minimum billing unit is 1 second afaik.
- serverlessmom 2y agoI think that’s flawed analysis, if you’re running FaaS then sure you can fail to see benefit from small improvements in time (AWS Lambda changed their billing resolution a few years back but before then the Go services didn’t save much money despite being faster) but if you’re running thousands of requests, and speeding them all up, you should be able to realize tangible compute savings whatever your platform.
- hiatus 2y agoHelp me to understand, then. If this stuff is being done on an autoscaling cluster, I can see it, but if you are just running everything on an always-on box for instance, it is less clear to me. edit: Do you have an affiliation with the blog? I ask because you have submitted several articles from checkly in the past.
- brunoarueira 2y agoI agree, but this post looks like an advertisement about the service itself.
- case 2y agoIt’s literally on the company’s blog, which is partially about promoting the company’s service. What’s the issue with that? (Long time happy Checkly user here, the service is fantastic)
- brunoarueira 2y agoNot a problem, but the OP is questioning about the savings! I, for example, like to dive more on insights like the relative savings vs absolut to learn the approaches other engineers take! It's all about metrics we should take care. (I'll put this service on my list to try someday, looks like fantastic indeed)
- redman25 2y agoI often ask myself the same question. We have some user facing queries that slow the frontend down. I’ve fixed some slowness but it’s definitely not a priority. I wonder how much speed improvements correlate with increased revenue by happy customers.
- encoderer 2y agoThink of this like changing the oil in your car. Over-optimizing is not going to help you at all but if you ignore it eventually it will all seize up. You have to keep that stuff in check.
- drakir_nosslin 2y agoBit late to the party, but companies report that webpage speed correlates with conversion. See e.g. https://www.cloudflare.com/en-gb/learning/performance/why-site-speed-matters/ https://www.cloudflare.com/en-gb/learning/performance/why-si... & https://www.cloudflare.com/en-gb/learning/performance/more/website-performance-conversion-rates/ https://www.cloudflare.com/en-gb/learning/performance/more/w... This one is also interesting; written in 2012, it claims that Amazon could lose 1b+ from a 1 sec slowdown: https://www.fastcompany.com/1825005/how-one-second-could-cost-amazon-16-billion-sales https://www.fastcompany.com/1825005/how-one-second-could-cos.... I imagine people are even less tolerant of slow pages today. Fixing website performance can be one of the cheapest ways to increase conversion because it's hard to figure out what else moves the needle.
- deleted 2y ago[deleted]
- dmurray 2y agoThe units seem wrong in any case. It's 3 months of compute per day, which is actually much more impressive. If we think about the business impact, we don't usually think of compute expenditure per-day, so you might reasonably say, the fix saved 90 years of annual compute. Looks better in your promotion packet, too.
- cdelsolar 2y agoμs isn't picoseconds, it's microseconds, which are a million times bigger...
- janOsch 2y agoThank you for pointing that out! You are correct, μs stands for microseconds, not picoseconds. I've corrected the mistake, and the update should be visible as soon as the CDN cache invalidates.
- serverlessmom 2y agoEvery day I have more sympathy for the Mars Climate Orbiter team. https://science.nasa.gov/mission/mars-climate-orbiter/ https://science.nasa.gov/mission/mars-climate-orbiter/
- harisamin 2y agoOn the noisy NodeJS auto-instrumentation, it is indeed very noisy out of the box. Myself along with a bunch of other ppl finally got the project to allow you to select the instrumentations via configuration. Saves having to create your own tracer.ts/js file. Here's the PR that got merged earlier in the year: https://github.com/open-telemetry/opentelemetry-js-contrib/pull/1953 https://github.com/open-telemetry/opentelemetry-js-contrib/p... The env var config is `OTEL_NODE_ENABLED_INSTRUMENTATIONS` Anyways, love Opentelemetry success stories. Been working hard on it at my current company and yielding fruits already :)
- roboben 2y agoThis is great. Last time I tried this I couldn’t even find a way in code to disable some.
- tnolet 2y agoThat is awesome. Had no idea this was available as an env var. After diving into OTel for our backend, we also found some of this stuff is just too noisy. We switch it of using this code snippet, for anyone bumping into this thread: instrumentations: [getNodeAutoInstrumentations({ '@opentelemetry/instrumentation-fs': { enabled: false, }, '@opentelemetry/instrumentation-net': { enabled: false, }, '@opentelemetry/instrumentation-dns': { enabled: false, },
- harisamin 2y agoyeah I totally turned all of those off...way too noisy :)
- serverlessmom 2y agoThis is so cool! I’ve had this exact problem before.
- Veserv 2y agoWhy would you disable instrumentation instead of just filtering the recorded log? That only makes sense if the instrumentation overhead itself is significant. But, for a efficient recording implementation that should only really start being a problem when your average span is ~1 us.
- bushbaba 2y agoJust a friendly call out that checkly is an awesome service.
- AtlasBarfed 2y agoIt's AWS, if you shake a stick at some network transfer optimization or storage/EBS/S3 you'll save three engineers salary.
- tnolet 2y agoThis deserves an updoot. We reached a level of scale right now at Checkly that all of these things start adding up. We moved workloads off of S3 to Cloudflare R2 because of this.
- tmpz22 2y ago> We moved workloads off of S3 to Cloudflare R2 because of this. So you moved from a mature but expensive storage solution to a younger currently subsidized storage solution? What happens when R2 jacks up pricing?
- tnolet 2y agolet me nuance that a bit. 99% of our workload is write heavy and is still on S3. We run monitoring checks that snap a screenshot and record a video. We write that to S3. Most folks will never view any of that as most checks pass and these artefacts only become interesting when things fail. Enter a new product feature we launched (Visual Regression Testing) which requires us to fetch an image from storage on every "run" we do. These could be every 10sec. This is where R2 shines. No egress cost for us. It's been rock solid and saved us about 60x compared to AWS. Still, we run most of our infra on AWS.
- 0xFF0123 2y ago> currently subsidized storage solution Interesting, do you have a source on the subsidized nature of R2?
- tmpz22 2y agoI do not and I'm likely misusing the word subsidized. My concern is that as a newer product (R2 in 2022 [1] compared to S3 2006 [2]) R2 has deliberately priced itself to compete with egress pricing of S3 in order to gain market share and developer mindshare. I am not confident Cloudflare will maintain this competitive pricing indefinitely as I expect it to follow well established industry trends of jacking up prices once a walled garden has been sufficiently establed. Further its my opinion cloud costs have grown at an absurd level as engineers and executives made poor and frankly lazy technology choices for the last decade. Ultimately I like cloudflare a lot but I think we need more discipline and lower operational overhead if we want infrastructure development to remain practical to individuals and small businesses versus mega-corps. Cloudflare with its free pricing tiers is often a default choice for these organizational sizes but it should not be viewed as a panacea and carries tradeoffs as with everything in life. [1] https://www.cloudflare.com/press-releases/2022/cloudflare-makes-r2-storage-available-to-all/ https://www.cloudflare.com/press-releases/2022/cloudflare-ma... [2] https://hidekazu-konishi.com/entry/aws_history_and_timeline_amazon_s3.html https://hidekazu-konishi.com/entry/aws_history_and_timeline_...
- philsnow 2y agoIs latency the same thing as duration? I think of latency as being more like a vector-with-starting-point (a “ray segment”?) than a scalar, it’s “rooted” to a point in time, so it doesn’t make sense to sum them.
- Cthulhu_ 2y agoGiven how high frequent this thing is, I'd say it's worth exploring moving away from Node; I don't associate Node with high performance / throughput myself.