4 ms·
> I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged There's a third perspective he
by ethbr1 1mo ago
> I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
There's a third perspective here: models are getting less useful, but overfitting to seeming useful to humans.
Imho, this is why analysis like TFA + third party cross-compatible harnesses (read: last mile UX) are so important to the leading labs optimizing for actual utility.
I'm suspicious enough of my subjective evaluation to believe a well-designed harness / verbiage could gaslight me into believing an objectively inferior model was superior. And at some point frontier labs are looking at the ROI of investing $1 in that vs actual model improvement.