3 ms·
> I’d suspect the harness to massively affect token use and optimisation Yes they do according to databricks -> https://www.databricks.com/blog/benchmarking-co
by maxignol 2mo ago
> I’d suspect the harness to massively affect token use and optimisation
Yes they do according to databricks -> https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase https://www.databricks.com/blog/benchmarking-coding-agents-d...
That's why I find comparing models on benchmarks only gives the tendency. We should be comparing model x harness to have accurate metrics.
- makingstuffs 2mo agoAgreed. An experiment is only as accurate as the methodology is sound. Testing each lab’s model in its own harness just tells us how well the lab has performed. For the model’s performance we need a control and the only way to get the control is to either test all models in all harnesses or all models in the same harness