5 ms·
I love OpenTelemetry and we want to trace almost every span happening. We’d be bankrupt if we went any vendor. We wired opentelemetry with Java magic, 0 effort
by CSDude 3y ago
I love OpenTelemetry and we want to trace almost every span happening. We’d be bankrupt if we went any vendor. We wired opentelemetry with Java magic, 0 effort and pointed to a self hosted Clickhouseand store 700m+ span per day with a 100$ EC2.
https://clickhouse.com/blog/how-we-used-clickhouse-to-store-opentelemetry-traces https://clickhouse.com/blog/how-we-used-clickhouse-to-store-...
- _boffin_ 3y agoAre you making sure that you're doing a sample rate, but send over all errors? At a former place, we were doing 5% of non-error traces.
- grogenaut 3y agoCareful, we've had systems go down under increased load just emitting errors if they didn't emit much in non error state
- _boffin_ 3y agoCan you go into more detail about your comment, please.
- cwp 3y agoNot the GP, but: Imagine you're sampling successful traces at, say, 1%, but sending all error traces. If your error rate is low, maybe also 1%, your trace volume will be about 2% of your overall request volume. Then you push an update that introduces a bug and now all requests fail with an error, and all those traces get sampled. Your trace volume just increased 50x, and your infrastructure may not be prepared for that.
- Longwelwind 3y agoI think what they means is that if you provisioned your system to receive spans for 5% of non-error requests and a few error requests, if for some random act of god, all the requests yield an error, your span collector will suddenyl receive spans for all requests.
- Topgamer7 3y agoWe've seen problems with memory usage on failure too. Python implementation sends data to the collector in a separate thread from the http server operations. But if these start failing, its configured for exponential backoff, so it can hold onto a lot of memory, and start causing issues with container memory limits.
- grogenaut 3y agoI've configured our systems to start dropping data at this point and emit an alarm metric that logging/metrics are overloaded
- grogenaut 3y agoSorry been busy running around all day. Basically what's happened for us on some very high transaction per second services is that we only log errors. Or Trace errors. And the service basically never has errors. So imagine a service that is getting 800,000 to 3 million request a second. And this is happily going along basically not logging or tracing anything. Then all the sudden a circuit opens on redis and for every single one of those requests that was meant to use that open circuit to redis you log or trace an error. You went from a system that is doing basically no logging or tracing to one that is logging or tracing at 800,000 to 3 million times a second. What actually happens is you open the circuit on redis because red is a little bit slow or you're a little bit slow calling redis and now you're logging or tracing 100,000 times a second instead of zero and that bit of logging makes the rest of the requests slow down and now you're actually within a few seconds logging or tracing 3 million requests a second. You have now toppled your tracing system your logging system and the service that's doing the work. Death spiral ensues. Now the systems that are calling this system starts slowing down and start tracing or logging more because they're also only tracing or logging mainly on error. Or sadly you have a better code that assumes that the tracing are logging system is up always and that starts failing causing errors and you get into doing extra special death loop that can only be recovered from by only attempting to log or error during an outage like this and you must push to fix. All the scenarios have happened to me in production. In general you don't want your system to do more work in a bad state. In fact as the AWS well architected guide say when you're overloaded or you're in a heavy air State you should be doing as little work as possible. So that you can recover
- CSDude 3y agoNormally yes, but we do a lot of data collection and identifying what's an error is usually hard because of partial errors. We also care about performance, per tenant and per resource with lots of dimensionality and sampling reduces that information for us.
- fastest963 3y agoHow do you send all errors? The way tracing works, as I understand it, is that each microservice gets a trace header which indicates if it should sample and each microservice itself records traces. If microservice A calls microservice B and B returns successfully but then A ends up erroring, how can you retroactively tell B to record the trace that it already finished making and threw away? Or do you just accept incomplete traces when there are errors?
- phamilton 3y agohttps://opentelemetry.io/docs/concepts/sampling/ https://opentelemetry.io/docs/concepts/sampling/ describes it as Head/Tail sampling, but in practice with vendors I see it as Ingestion sampling and Index sampling. We send all our spans to be ingested, but have a sample rate on indexing. That allows us to override the sampling at index and force errors and other high value spans to always be indexed.
- fastest963 3y agoMaybe the Go client doesn't support that? https://opentelemetry.io/docs/instrumentation/go/sampling/ https://opentelemetry.io/docs/instrumentation/go/sampling/
- phillipcarter 3y agoIt does, but the docs aren't clear on that yet. TraceIdRatioBased is the "take X% of traces" sampler that all SDKs support today.
- sigwinch28 3y agoYou can do head-based sampling and tail-based sampling. With head sampling, the first service in the request chain can make the decision about whether to trace, which can reduce tracing overhead on services further down. With tail-based sampling, the tracing backend can make a determination about whether to persist the trace after the trace has been collected. This has tracing overheads, but allows you to make decisions like “always keep errors”.
- push-to-prod 3y agoThat's a really informative post, the ClickHouse thing sounds interesting!
- annanay 3y agoThis is really interesting, thanks for sharing. What's also cool was the low effort needed for this setup (Java autoinstrumentation + Clickhouse exporter + Grafana Clickhouse Plugin).
- podoman 3y agoThe reality is that most people don't want to manage their own Clickhouse store, and not all engineers can operate with SQL as efficiently as with code (me included). Nonetheless, this is pretty cool!
- klysm 3y agoSQL is code and absolutely worth learning.
- simonw 3y agoI'm beginning to sound like a broken record at this point, but if you don't know SQL very well but know how to use GPT-4, you have access to enough SQL to get a lot more done than you might think.
- hnarn 3y ago> not all engineers can operate with SQL as efficiently as with code I don’t mean for this to sound insulting but I honestly do not think this is an acceptable take to have as a developer. Not knowing SQL is like refusing to learn any language that has classes in it, simply because you don’t like it. I’ve heard stories of huge corporations failing product launches because some code was written to SELECT * from a database and filtering it in-app instead of doing the queries correctly, and what’s so fun with these types of issues is that they usually don’t appear until weeks later when the table has grown to a size where it becomes a problem. When you’re saying that you’d rather find the data in-app than in-database, you’re putting the work on an inferior party in the transaction simply because you can’t be bothered. The code will never* find the correct data faster than the database. * there may be exceptions, but they’re far enough between to still say “never”.
- xyzzy_plugh 3y agoDropping down to SQL to write a really complex query is, in my professional experience, always a poor use of time. It's far simpler to just write the dumb for-loops over your data, if you can access it. Of course not all engineers can operate with SQL as efficiently as code -- that's the whole point. Otherwise why would we be writing code? Learning SQL intimately doesn't change that fact.
- flaviut 3y agoI've got a small personal project submitting traces/logs/metrics to Clickhouse via SigNoz. Only about 400k-800k spans per day (https://i.imgur.com/s0J6Mzo.png https://i.imgur.com/s0J6Mzo.png), but running on a single t4g.small with CPU typically at 11% and IOPS at 4%. I also have everything older than a certain number of GB getting pushed to a sc1 cold storage drive. w/ 1 month retention for traces: ┌─parts.table─────────────────┬──────rows─┬─disk_size──┬─engine────┬─compressed_size─┬─uncompressed_size─┬────ratio─┐ │ signoz_index_v2 │ 26902115 │ 17.06 GiB │ MergeTree │ 6.21 GiB │ 66.74 GiB │ 0.0930 │ │ durationSort │ 26901998 │ 5.44 GiB │ MergeTree │ 5.40 GiB │ 53.02 GiB │ 0.10190 │ │ trace_log │ 123185362 │ 2.64 GiB │ MergeTree │ 2.64 GiB │ 37.96 GiB │ 0.0695 │ │ trace_log_0 │ 120052084 │ 2.46 GiB │ MergeTree │ 2.45 GiB │ 37.60 GiB │ 0.06528 │ │ signoz_spans │ 26902115 │ 2.21 GiB │ MergeTree │ 2.21 GiB │ 76.73 GiB │ 0.028784 │ │ query_log │ 16384865 │ 1.91 GiB │ MergeTree │ 1.90 GiB │ 18.31 GiB │ 0.10398 │ │ part_log │ 17906105 │ 846.73 MiB │ MergeTree │ 845.39 MiB │ 3.84 GiB │ 0.21521 │ │ metric_log │ 4713151 │ 820.92 MiB │ MergeTree │ 806.13 MiB │ 14.56 GiB │ 0.05405 │ │ part_log_0 │ 15632289 │ 702.82 MiB │ MergeTree │ 701.70 MiB │ 3.34 GiB │ 0.20490 │ │ asynchronous_metric_log │ 795170674 │ 576.24 MiB │ MergeTree │ 562.50 MiB │ 11.11 GiB │ 0.049429 │ │ query_views_log │ 6597156 │ 461.35 MiB │ MergeTree │ 459.75 MiB │ 6.36 GiB │ 0.07060 │ │ logs │ 6448259 │ 408.59 MiB │ MergeTree │ 406.65 MiB │ 5.99 GiB │ 0.06627 │ │ samples_v2 │ 949110122 │ 345.01 MiB │ MergeTree │ 325.31 MiB │ 22.09 GiB │ 0.014382 │ If I was less stupid I'd get a machine with the recommended Clickhouse specs and save myself a few hours of tuning, but this works great. Downsides: - clickhouse takes about 5 minute to start up because my tiny sc1 drive has like 4 IOPS allowed - signoz's UI isn't amazing. It's totally functional, and they've been improving very quickly, but don't expect datadog-level polish
- pranay01 3y agoThanks for mentioning SigNoz, I am one of the maintainers at SigNoz and would love your feedback on how we can improve it further. If anyone wants to check our project, here’s our GitHub repo - https://github.com/SigNoz/signoz https://github.com/SigNoz/signoz
- deleted 3y ago[deleted]