6 ms·
Good question! Depends on the instrumentation libraries for sure, but in our Go services we've been using Jaeger protocol for our traces, and "sampling" at 100%
by strofcon 6y ago
Good question!
Depends on the instrumentation libraries for sure, but in our Go services we've been using Jaeger protocol for our traces, and "sampling" at 100%, with negligible impact to request / response times.
We spit out some 30k+ spans per second, FWIW. :-)
Edit: Disclaimer, we're not using Tempo.
- sna1l 6y agoThanks for the reply! I think I must be missing something, but it seems like the big difference between Tempo and a traditional tracing system is the storage indexing & database (ES/C* vs object store and index all fields vs key/value lookup by ID). I vaguely remember reading something that latency even from EC2 -> S3 can be around 200-300ms. Wouldn't this cause the overhead to rise? Feel free to point me to any documentation that clear this up!
- netingle 6y agoWrites are batches up and committed asynchronously to s3 - this should add much if any latency to your services.
- number101010 6y agoWe are ingesting 170k+ with Tempo. It is 100% of our read/query path. Disclaimer: I am using Tempo :) (and from Grafana)
- RhodesianHunter 6y agoEven with batching from Tempo, wouldn't that cost many thousands per month in S3 PUT costs alone?
- number101010 6y agoWe batch up traces in a block and write a block at a time. Internally we are currently configured to write 100k traces in one batch.
- atombender 6y agoDoesn't this cause explosive memory usage? What happens if there's some congestion? Is there a circuit breaker to start dumping (discarding) log entries past a certain limit? I was testing Google Pub/Sub's Go client for publishing internal API event data for later ingest to BigQuery, and it turns out Pub/Sub publishing is not that much faster than writing directly to BigQuery. The buffer sizes we'd need to avoid adding latency to our APIs would have to be ridiculously high; the Pub/Sub client buffers and submits batches in the background (its default buffer size is 100MB!). I don't like the idea of having huge buffers that increase with the request rate. Conversely, pushing the data to NATS in recent time without any buffering or batching turned out to be fast enough to not add any latency. You have to be able to receive messages very fast on the consumer side (as NATS will start dropping messages if consumers can't keep up), but you can simply run a few big horizontally autoscaled ingest processes that can sit there ingesting as fast as they can, which never impacts API latency at all.
- marcinzm 6y agoS3 PUTs are $0.005 per 1000. If you're writing twice a second that comes out to $25/month.
- RhodesianHunter 6y agoYeah, but the person I'm responding to is suggesting 170k+ spans per second so how is twice per second relevant?