3 ms·
Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run no
by claudiusa 2mo ago
Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of the table.
- IreneAI 2mo ago[flagged]