3 ms·
Your comment brings up a good point (and also one of our big challenges): there is a huge diversity in the tools teams use to setup and operate their infra. Rig
by wilson090 2y ago
Your comment brings up a good point (and also one of our big challenges): there is a huge diversity in the tools teams use to setup and operate their infra. Right now our platform only speaks to your cluster directly through kubectl commands. We’ll build other integrations so it can communicate with things like Elastic Search to broaden its context as needed, but we’ll have to be somewhat thoughtful in picking the highest ROI integrations to build.
Currently, we only handle the investigation piece and suggest a remediation to the on-call engineer. But to properly move into automatically applying a fix, which we hope to do at some point, we’ll need to integrate into CI/CD
As for the demo example, I agree that the issue itself isn’t the most compelling. We used it as an example since it is easy to visualize and set up for a demo. The agent is capable of investigating more complex issues we've seen in our customer's production clusters, but we're still looking for a way to better simulate these on our test environment, so if you/anyone has ideas we’d love to hear them.
We do think this has more value for engineers/teams with less expertise in k8s, but we think SREs will still find it useful
- stackskipton 2y ago>we're still looking for a way to better simulate these on our test environment, so if you/anyone has ideas we’d love to hear them. Pick Kubernetes offering from big 3, deploy it then blow it up. (I couldn't get HackerNews to format properly and done fighting it) On Azure, deploy a Kubernetes cluster with following: Azure CNI with Network Policies Application Gateway for Containers External DNS hooked to Azure DNS Ingress Nginx Flexible PostGres Server (outside the cluster) FluxCD/Argo Something with using Workload Identity Once all that is configured, put some fake workloads on it and start misconfiguring it with your LLM wired up. When the fireworks start, identify the failures and train your LLM properly.
- solatic 2y ago> we think SREs will still find it useful There are two kinds of outages: people being idiots and legit hard-to-track-down bugs. SREs worth their salt don't need help with the former. They may find an AI bot somewhat useful to find root cause quicker, but usually not so valuable as to justify paying the kind of price you would need to charge to make your business viable to VCs. As for the latter, good luck collecting enough training data. Otherwise, you're selling a self-driving car to executives who want the chauffeur without the salary. Sounds like a great idea, until you think about the tail cases. Then you wish you had a chauffeur (or picked up driving skills yourself). Maybe you'll find a market, but as an SRE, I wouldn't want to sell it.