4 ms·
> Performance is significantly higher than Fable 5.1 That's not clear. Need to see independent benchmarks first.
by andxor 1mo ago
> Performance is significantly higher than Fable 5.1
That's not clear. Need to see independent benchmarks first.
- andxor 1mo agoArtificial Analysis just published their aggregate score (61). Still below Fable 5, let alone Fable 5.1. EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
- timpera 1mo agoI agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
- CamperBob2 1mo agoThere is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy. If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
- natsucks 1mo agoI saw this too and I'm really confused.
- forgot-my-pw 1mo agoWe need them pelicans on bikes.
- bwat49 1mo agoIts time to move on to the flamingo on a unicycle bench
- forgot-my-pw 1mo agoAA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra https://artificialanalysis.ai/articles/benchmarking-gpt-6-as... TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
- tintor 1mo ago66 vs 61 is not 'about the same'. GPT 5.6 is also 61 like Astra.
- forgot-my-pw 20d agoLate reply: I'm not sure which chart you're referring to. The top 2 charts shows Astra has the same score as Fable 5.1. Even the title says "GPT-6 Astra ties leadership with Claude Fable 5.1 in both of our flagship Indices, at lower cost. Astra equals Fable 5.1 in the Intelligence Index at ~40% of the cost, and in the Coding Agent Index at ~60% of the cost."