3 ms·
Artificial Analysis just published their aggregate score (61). Still below Fable 5, let alone Fable 5.1. EDIT: This is suspiciously low. Calls the relevance o
by andxor 1mo ago
Artificial Analysis just published their aggregate score (61).
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
- timpera 1mo agoI agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
- CamperBob2 1mo agoThere is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy. If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
- natsucks 1mo agoI saw this too and I'm really confused.