4 ms·
we have a lot of guardrails (https://docs.hyperprobe.co/how-it-works#built-in-safety-guardrails https://docs.hyperprobe.co/how-it-works#built-in-safety-guar...)
by karanraina 2mo ago
we have a lot of guardrails (https://docs.hyperprobe.co/how-it-works#built-in-safety-guardrails https://docs.hyperprobe.co/how-it-works#built-in-safety-guar...)
if any guardrails fails, we suspend probes till cooldown.
also
1. every probe is bounded by hits/expiry time (whichever comes earlier)
2. hit budgeting happens with a token bucket at a global level, per probe was an overkill (numbers are configurable)
3. we even measure the execution time that probes have when active and suspend if that that takes longer than threshold (again configurable)
4. we even have budgets for the network bandwidth it would take (approximated by the size of payloads)
5. collection itself is bounded by max no of total snapshots we can keep in memory.
6. every snapshot has a size limit as well, every variable has a size limit as well.
7. depth of objects, no of objects, size of lists is capped by default.
latency delta varies by platform under load but is mostly negligible
nodejs: ~7-10ms
python: ~4-9ms
java: 1-2ms
the main reason for this is guardrails suspending probes, having loosened guardrails will increase this under load
regarding localization of failures.. absolutely
we even report the error in the probe snapshot (confirmed by adding side effects in an expression and commenting out the guardrails during testing)
huge payload size doesnt matter.. we limit the objects depth, list length, remove duplicate refs from data etc.. even string length is truncated., but even if it happens, your request would still survive.
also, even if the collector dies or there's a network failure, your service remains unaffected, we just are unable to collect telemetry
we are boring under extreme conditions :)