4 ms·
And tbis is why we have logging system. A decent automated query of your syslog output should be able to show you any flapping services. A setup like this giv
by SkipperCat 3y ago
And tbis is why we have logging system. A decent automated query of your syslog output should be able to show you any flapping services. A setup like this gives you not only self-healing (as described in the post) but also visibility into erroneous systemd units.
I always thought a good monitoring system had three components. Polling, streaming data and log analysis. A lot of times folks don't bother with the logging and miss issues like what is described in the post.
- spmurrayzzz 3y agoYea I was a little surprised to see the comment "It wouldn't hurt to look at the logs or the metrics, just in case" paired with the previous comment "there's not quite enough information exposed in the Prometheus host agent's systemd metrics to make it easy". Not everything needs to be monitored in a single way via a single platform. A cron job that greps/awks syslog for relevant strings and sends a slack message, email, sms, something would be a step function in the right direction.
- yabones 3y agoIt's even easier than that, the functionality exists inside rsyslog. All you have to do is set up a template in /etc/rsyslog.d and point it towards a mail relay... if $syslogseverity <= 3 then { action(type="ommail" server="127.0.0.1" port="25" mailfrom="rsyslog@localhost" mailto="root@localhost" subject.template="mailSubject" template="mailBody" action.execonlyonceeveryinterval="3600") } While I generally like to dump all logs into elasticsearch/opensearch, it's also really handy to have this set up for one-off machines so that critical/error logs that usually come from things breaking & restarting still get seen. https://docbot.onetwoseven.one/services/syslog/ https://docbot.onetwoseven.one/services/syslog/
- ilyt 3y agoRsyslog also supports sending to elasticsearch IIRC.
- ilyt 3y agoSemi-related but it's shame default cron in most (all?) kinda fucking sucks at everything. There is no builtin spreading so */5 job on 500 nodes will get you nasty spike in CPU load. Logging options are pretty much "parse emails and hope for best" and some syslog spam that's mostly noise if everything is working fine. And fucked up environment vars that trip many script-makers. No metrics whatsoever
- spmurrayzzz 3y agoOh definitely. This actually bit me a few years ago. Had a distributed set of embedded systems (tens of thousands of wifi routers) that were phoning home which I didn't realize was being done through a cron. It was something we inherited with the BSP. Looking at our analytics for the relevant API zoomed out to 1-5 minute resolutions, our RTTs looked like they were monotonically increasing over time. We had to hookup real time analytics to one of the nodes in our production cluster to be able to catch the spikes at the exact second of the relevant minute boundary. The solution in our case was to build a tiny launcher for anything cron-related that delayed the script executions randomly within a range of milliseconds.