3 ms·
But your p99 for endpoint B hasn’t changed, why would you alert based on service-wide stats? It’s easy to tag/dimension metrics by urls
by falsedan 7y ago
But your p99 for endpoint B hasn’t changed, why would you alert based on service-wide stats? It’s easy to tag/dimension metrics by urls
- viraptor 7y agoChange the level of caching and you get the same result. Let's say you have the same endpoint, but start caching user X but not others, because reasons.
- falsedan 7y agoStill struggling, if requests from X were slow, caching their responses would only drop or maintain p99. If those requests were fast, p99 might increase & alert... but if you were caring about performance at a user-level, you can imagine you're dimensioning on users too, and you'd see the p99s for each user hadn't changed, and would shrug and bump the alert threshold up. If it were me, I'd probably start by caching the slow requests (that are above the p99) first & adjusting the alert thresholds to match the new performance.