4 ms·
A series of hot takes: Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model perfor
by albrewer 1mo ago
A series of hot takes:
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
- jdm2212 1mo agoI think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.