3 ms·
Congrats on the launch. Having worked at a startup in the AIOps space, I can can offer a few suggestions. 1. No matter how good your AI, it will make mistakes.
by csears 6y ago
Congrats on the launch. Having worked at a startup in the AIOps space, I can can offer a few suggestions.
1. No matter how good your AI, it will make mistakes. Users need ways to provide feedback or filtering to avoid bad alert fatigue. Giving users a sense of control is critical.
2. Most larger shops will have dozens of monitoring tools already generating alerts. Consider ingesting existing alerts as another algorithmic signal.
3. The real root cause of an incident often won't show up in logs. Don't assume that the earliest event in a cluster is causal.
4. The more context you can provide an operator looking at a potential incident, the better. Modern SEIM tools do an ok job here. Consider pulling in topology or other enrichment sources and matching entity names/IDs to log data.
Good luck. Contact in profile if you'd like to chat further.
- Ajs1 6y agocsears, founder here - could not agree more with your comments. 1. we learnt early to make user feedback easy (and immediately actionable) - they can quickly "like", "mute" and "spam". Or go more granular if needed. 2. This is an insightful comment. A few of our early users gave us similar feedback, and we've been hard at work. We'll soon be releasing a mode that takes an incident signal from your incident management tool such as PagerDuty or even Slack (often people create a Slack workspace per incident), and constructs a report around it. 3&4 are good points as well. Don't disagree about enrichment, just need to stage things.