2 ms·
Evaluating performance and efficiency of the GitHub Copilot agentic harness
- brammertottens 4mo agoIt's an interesting post, but i'm a bit skeptical on their decision to report the best run for each agent, and not just the mean over the 5 runs. We have seen this as well in running benchmarks, that variance within one setup can be pretty big.