4 ms·
Thanks for the great article! At my current employer we have a company-wide service for aggregating error logs in particular (WARN, ERORR level log rows and st
by ekiauhce 4y ago
Thanks for the great article!
At my current employer we have a company-wide service for aggregating error logs in particular (WARN, ERORR level log rows and stacktraces, if it was an exception) so developers can analyze them for debugging purposes. Also it automatically gathers information about incoming http request (geo, ip address, user agent, etc) and you can easily see a particular segment of errors, and what kind of users getting them.
As I can see you have logs quantitative metric https://community-demo.coroot.com/p/oc1vhnmq/app/default:Deployment:cart/Logs https://community-demo.coroot.com/p/oc1vhnmq/app/default:Dep... but without any detalization (maybe it works this way only for the demo app). I mean, it would be great to be able to inspect each ERROR event separately or to define custom SLO with alert for particular type of errors.
Another great feature we use a lot is historical data, so you can find patterns of error spikes on months scale and when it has gone after fix.
FYI this error-service I'm talking about is built on top of the ClickHouse, so it's quite responsive regardless of the large volumes of data.
Another thing I want to mention is cron-like workload (or batch jobs, you name it). Is there any support or useful metrics for it?