4 ms·
How is HyperProbe different from existing tools like AppSignal, Rollbar, and Embrace? Such very mature tools exist that auto-instrument, collect variables from
by doublerebel 2mo ago
How is HyperProbe different from existing tools like AppSignal, Rollbar, and Embrace? Such very mature tools exist that auto-instrument, collect variables from the call stack, and pinpoint error causes.
> Every log-and-trace tool hands the agent data that already exists and asks it to reason backward to what probably happened
If the app is using a decent instrumentation tool, the data shows what 'actually' happened, not what 'probably' happened.
> "checkout returns 200 but some users are seeing their order fail, find out why."
Does this tool only exist to shore up poor system design? Failing orders at any e-commerce business I've worked with, large and small, are a huge red flag. Typically that is one of the first actions that is logged and traced (alongside onboarding/login), and the metrics are actively monitored. Returning 200 for failure and not catching that error is very bad API design.
Similarly, putting engineers in a situation where debugging requires accessing unknown amounts of live sensitive customer data is generally considered bad practice (even if it happens often IRL) -- in a hurry to debug, it's easy to miss that a property should have been redacted; by then it's too late and sensitive data is exposed. Plus, in most systems with significant usage the volume of trace data is prohibitive to individually examine and search through. That's why Rollbar etc aggregate errors and captured data to identify patterns before a human (or agent, or tool) ever takes a look at it. A single captured instance can also be very misleading as to the true cause.
How are you addressing these common concerns?
- karanraina 2mo ago> How is HyperProbe different from existing tools like AppSignal, Rollbar, and Embrace? These work only on either uncaught exceptions or wrapping up caught exceptions with their sdk. These tools will not help you with silent failures, like logic bugs where code executes cleanly without throwing, but produces the wrong business state. If every problem in your app ends up as an exception, sure you'll be able to catch the symptoms of where the exception got thrown. we can deal with these too, but these tools cant deal with the messy bugs where no exception fires. > Such very mature tools exist that auto-instrument, collect variables from the call stack, and pinpoint error causes. That is true for python using frame.f_locals (we use this as well) nodejs only gives it only till the lasy async boundary, after that v8 itself drops this data. java only gives you the current frame, to get variables beyond that you would needs JDI/JVMTI which would block your threads, usually unnacceptable in production To get around this safely, we add multiple probes all across the call chain and collate collected data using the traceId from the context (or thread id as a fallback); > Does this tool only exist to shore up poor system design? Returning 200 OK on a silent failure is 100% bad system design, I completely agree. But real-world production systems are full of legacy edge cases. (if that weren't true, L1/L2/L3 support team shenanigans wouldn't exist) Also, the exception will tell you that an exception occured in order service in GET /orders/{id}/payment, your trace will tell you payment service is giving 404 for that order ID what it wont tell you it happened becuase the webhook endpoint that your payment gateway calls is now receiving a new payment state called 'PENDING' and that you dont handle but still mark the payment as 'processed' for idempotency check. and now your order service is calling the payment service and its giving 404 because it never got written Bad design. 100% Agree, but has happened IRL. > putting engineers in a situation where debugging requires accessing unknown amounts of live sensitive customer data is generally considered bad practice (even if it happens often IRL) I think tells that teams would go to these extents to fix issues. Not ideal. I agree. > in a hurry to debug, it's easy to miss that a property should have been redacted; by then it's too late and sensitive data is exposed fair critique. we currently use in-process rule engines to filter known sensitive patterns, and users can add on to it. but we are also building out-of-process secondary checks (using NER/classifiers) to sanitize payloads before storage. It requires strict rules, but getting verified runtime evidence is far safer and faster than blindly guessing and shipping trial-and-error hotfixes to production. or waiting to be too sure.. a luxury that might not be possible everytime. > Rollbar etc aggregate errors and captured data to identify patterns before a human (or agent, or tool) ever takes a look at it. There is merit in that as well, if you are looking at so many logs/traces, you kinda have to do it. We have a different approach, we use hypothesis driven conditional probing instead. probes are dropped dynamically as the understanding of the bug evolves in a session exmaple: console.log('hello'); const x = await getThisValueSomehow(); if (condition A) { console.log('i m in condition A'); // do something; } else if (condtion B) { console.log('i m in condition B'); // do something; } You can also place a probe before the branch to capture variable state when neither condition evaluates to true. You gather precise data on demand rather than paying to store petabytes of static trace data. > A single captured instance can also be very misleading as to the true cause. We collect multiple snapshots per probe run. However, because we capture full variable state at the exact execution line, a single snapshot frequently reveals the root cause for that specific failure path. If that snapshot raises new questions, you/your agent simply drops more probes deeper down the call chain Thanks! This was very insightful