4 ms·
Yeah, there's a big difference between something that's pure distributed tracing (this) and something that does metrics/monitoring as well (TraceView, New Relic
by dkuebric 10y ago
Yeah, there's a big difference between something that's pure distributed tracing (this) and something that does metrics/monitoring as well (TraceView, New Relic, ...). Sampling is a valid way to make distributed tracing scalable and performant, but it can also limit the use-cases for the data.
I imagine this being used for episodic debug cases, where one could turn up the sample rate and pay only for the traces captured/queried during an incident. But you would need to do your monitoring and trending separately.
I wouldn't be surprised if they eventually start integrating this better with CloudWatch for that reason, though it doesn't seem to be doing any of that today.
Disclosure: I work on TraceView (traceview.solarwinds.com) which is distributed tracing based APM product. We're inspired by Google Dapper and x-trace, both mentioned elsewhere in this thread.
- DenisM 10y ago> difference between something that's pure distributed tracing (this) and something that does metrics/monitoring as well (TraceView, New Relic, ...) I don't see the distinction you're making, as neither seems to record all traces for accurate audit. Am I missing it?
- dkuebric 10y agoI was speaking to a more general monitoring approach where you might want to know p99 latency, request volume, error rate, etc, for each service to use in alerting and trending. This is a common use-case for application monitoring that isn't addressed well by a pure tracing approach. If you're looking specifically for the 100% audit trail case, I'd look at DynaTrace or potentially Instana.