3 ms·
seems really interesting @kirtivr . Looking forward to it. Though, i wonder if it could also be an integration to alerting platforms directly (like NewRelic, D
by nxtcoder17 3mo ago
seems really interesting @kirtivr . Looking forward to it.
Though, i wonder if it could also be an integration to alerting platforms directly (like NewRelic, Datadog etc.), so that for on-call alerts, it could cover the foundation work, and have some hypotheses ready for the on-call engineer to directly jump into
- kirtivr 3mo agoThis is on our roadmap! At this time we are focussing on evaluation benchmarks like SWE-bench (verified). This is a simpler benchmark and does not really map well to investigating alerts that have a huge amount of context. But its a start. I am wondering if we can improve upon foundation models with our reproduction <-> hypothesis loop approach. Foundation models tend to be precision first, and context limited, so they can get sidetracked by various things. This is an interesting problem to be working on right now!