3 ms·
Two things I would want to know before pointing this at a hot service: (1) the overhead budget — when a probe lands on a hot path, is capture sampled or capped
by tizerluo 2mo ago
Two things I would want to know before pointing this at a hot service: (1) the overhead budget — when a probe lands on a hot path, is capture sampled or capped per hit, and what p99 latency delta have you measured under load? (2) failure isolation — if probe evaluation itself throws (weird object shape, getter with side effects, huge captured value to serialize), is it contained so it cannot take down the request it is observing? In-process agents live or die by staying boring under worst-case conditions.
- karanraina 2mo agowe have a lot of guardrails (https://docs.hyperprobe.co/how-it-works#built-in-safety-guardrails https://docs.hyperprobe.co/how-it-works#built-in-safety-guar...) if any guardrails fails, we suspend probes till cooldown. also 1. every probe is bounded by hits/expiry time (whichever comes earlier) 2. hit budgeting happens with a token bucket at a global level, per probe was an overkill (numbers are configurable) 3. we even measure the execution time that probes have when active and suspend if that that takes longer than threshold (again configurable) 4. we even have budgets for the network bandwidth it would take (approximated by the size of payloads) 5. collection itself is bounded by max no of total snapshots we can keep in memory. 6. every snapshot has a size limit as well, every variable has a size limit as well. 7. depth of objects, no of objects, size of lists is capped by default. latency delta varies by platform under load but is mostly negligible nodejs: ~7-10ms python: ~4-9ms java: 1-2ms the main reason for this is guardrails suspending probes, having loosened guardrails will increase this under load regarding localization of failures.. absolutely we even report the error in the probe snapshot (confirmed by adding side effects in an expression and commenting out the guardrails during testing) huge payload size doesnt matter.. we limit the objects depth, list length, remove duplicate refs from data etc.. even string length is truncated., but even if it happens, your request would still survive. also, even if the collector dies or there's a network failure, your service remains unaffected, we just are unable to collect telemetry we are boring under extreme conditions :)