3 ms·
> How is HyperProbe different from existing tools like AppSignal, Rollbar, and Embrace? These work only on either uncaught exceptions or wrapping up caught exc
by karanraina 2mo ago
> How is HyperProbe different from existing tools like AppSignal, Rollbar, and Embrace?
These work only on either uncaught exceptions or wrapping up caught exceptions with their sdk.
These tools will not help you with silent failures, like logic bugs where code executes cleanly without throwing, but produces the wrong business state. If every problem in your app ends up as an exception, sure you'll be able to catch the symptoms of where the exception got thrown. we can deal with these too, but these tools cant deal with the messy bugs where no exception fires.
> Such very mature tools exist that auto-instrument, collect variables from the call stack, and pinpoint error causes.
That is true for python using frame.f_locals (we use this as well)
nodejs only gives it only till the lasy async boundary, after that v8 itself drops this data.
java only gives you the current frame, to get variables beyond that you would needs JDI/JVMTI which would block your threads, usually unnacceptable in production
To get around this safely, we add multiple probes all across the call chain and collate collected data using the traceId from the context (or thread id as a fallback);
> Does this tool only exist to shore up poor system design?
Returning 200 OK on a silent failure is 100% bad system design, I completely agree. But real-world production systems are full of legacy edge cases. (if that weren't true, L1/L2/L3 support team shenanigans wouldn't exist)
Also, the exception will tell you that an exception occured in order service in GET /orders/{id}/payment, your trace will tell you payment service is giving 404 for that order ID
what it wont tell you it happened becuase the webhook endpoint that your payment gateway calls is now receiving a new payment state called 'PENDING' and that you dont handle but still mark the payment as 'processed' for idempotency check.
and now your order service is calling the payment service and its giving 404 because it never got written
Bad design. 100% Agree, but has happened IRL.
> putting engineers in a situation where debugging requires accessing unknown amounts of live sensitive customer data is generally considered bad practice (even if it happens often IRL)
I think tells that teams would go to these extents to fix issues. Not ideal. I agree.
> in a hurry to debug, it's easy to miss that a property should have been redacted; by then it's too late and sensitive data is exposed
fair critique. we currently use in-process rule engines to filter known sensitive patterns, and users can add on to it. but we are also building out-of-process secondary checks (using NER/classifiers) to sanitize payloads before storage. It requires strict rules, but getting verified runtime evidence is far safer and faster than blindly guessing and shipping trial-and-error hotfixes to production. or waiting to be too sure.. a luxury that might not be possible everytime.
> Rollbar etc aggregate errors and captured data to identify patterns before a human (or agent, or tool) ever takes a look at it.
There is merit in that as well, if you are looking at so many logs/traces, you kinda have to do it.
We have a different approach, we use hypothesis driven conditional probing instead. probes are dropped dynamically as the understanding of the bug evolves in a session
exmaple:
console.log('hello');
const x = await getThisValueSomehow();
if (condition A) {
console.log('i m in condition A');
// do something;
} else if (condtion B) {
console.log('i m in condition B');
// do something;
}
You can also place a probe before the branch to capture variable state when neither condition evaluates to true. You gather precise data on demand rather than paying to store petabytes of static trace data.
> A single captured instance can also be very misleading as to the true cause.
We collect multiple snapshots per probe run. However, because we capture full variable state at the exact execution line, a single snapshot frequently reveals the root cause for that specific failure path. If that snapshot raises new questions, you/your agent simply drops more probes deeper down the call chain
Thanks! This was very insightful