2 ms·
Yet, AI is not there yet. Even the top models struggle at simplest SRE tasks. We just created a benchmark on adding distributed logs (OpenTelemetry instrument
by stared 9mo ago
Yet, AI is not there yet. Even the top models struggle at simplest SRE tasks.
We just created a benchmark on adding distributed logs (OpenTelemetry instrumentation) to small services, around 300 lines of code.
Claude Opus 4.5 succeed at 29%, GPT 5.2 at 26%, Gemini 3 Pro at 16%.
https://quesma.com/blog/introducing-otel-bench/ https://quesma.com/blog/introducing-otel-bench/