3 ms·
Ops is maybe trying to help developers and increase security (no login on the boxes needed anymore to look at logs). Better workflow for (certain kind of) alert
by chronid 9y ago
Ops is maybe trying to help developers and increase security (no login on the boxes needed anymore to look at logs). Better workflow for (certain kind of) alerts, which may then get more complicated. Maybe kibana dashboarding. That in turn will help developers tracking logs for each requests (as someone mentioned above, with a common UUID passed around).
And in there lies their mistake: provide the functionality (single access point to logs, alerting on logs) and make the developers work to gain access to it like it was another service. You want ES access? Parse your logs, I'm not going to write filters for you (I can help, but you are going to do the bulk of the work).
This guarantees logs in a parsable format in less than 3 months of developers being oncall, particularly if you have a container scheduler (we're talking about microservices right?) and use ASGs (and your policy for upgrades and box issues is "blow up the box and create another").
- flukus 9y ago> Ops is maybe trying to help developers and increase security (no login on the boxes needed anymore to look at logs). Can this be handled fine by having tail run as a daemon and forwarding all log entries to somewhere that the developers do have access? > Better workflow for (certain kind of) alerts, which may then get more complicated. This is likely to be highly specific as to what sort of alerts you want, but will tailing the log and running it through an awk script work? I assume you'll have to do something much like this with any alert management tool. Tools like kabana like nice, but this is one of those areas where I think people might be going for the shiny solution (that looks great to management) instead of analyzing what they really need.
- chronid 9y agoYou can use whatever system you want (syslog can drop the logs in a single box, no need for horrible tail-based concoctions), as long as you mantain it and it's not Ops responsibility to fix your AWK scripts that look for events in logs from 3 different services for the same request and the same customer (and/or respond to the alerts those script generate at 2am after Bob forgot to update it to parse correctly a new log message - we have our own fires to fight already). I'm not saying one should go straight to ELK. There are other ways, but at the end of the day you are going to implement a similar stack, guaranteed, and you are going to regret using freeform logging instead of a sensible structured format.
- flukus 9y agoIt sounds like Ops want all the control but to do none of the work? Maintaining scripts like this is there job.
- chronid 9y agoI'm sorry: it's your application, your alerts, your responsibility. Not Ops.
- flukus 9y agoIf it's my responsibility then I need access to the systems, otherwise I can't perform excercise that responsibility. How does alert creation and maintenance not fall under ops? They sound like the sort of ops people that devops teams were created to replace.
- JdeBP 9y agoForwarding, yes. tail, no. tail will lose log data. * http://jdebp.eu./FGA/do-not-use-logrotate.html#Problems http://jdebp.eu./FGA/do-not-use-logrotate.html#Problems