3 ms·
I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is
by gizmodo59 1mo ago
I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The fact that muse and 3.8 were high a day back shows they are just the modern version of lmareana for the mass audience and PR stunts.
- jesuslop 1mo agoIs there something better over there you'd recommend?
- yorwba 1mo agoThe Epoch Capabilities Index uses an Elo-based aggregation method that dynamically adjusts for benchmark difficulty and they put error bars on their scores, both of which put them miles ahead of Artificial Analysis: https://epoch.ai/eci?view=graph&tab=leaderboard https://epoch.ai/eci?view=graph&tab=leaderboard
- testycool 1mo agoAlso ArtificialAnalysis drop older foundational models to make room for new ones. If you customize the filter to add Gemini 3.1 Pro you'll see it ranks 5th in AA-Omniscince Index. Yet by default you wont see Gemini 3.1 Pro. I find this to be very unhelpful and confusing.
- jascha_eng 1mo agoBut it is actually a great model it e.g. got the carwash question right from 9 months ago. While openais models all struggled.
- jascha_eng 1mo agoI dont think you read my message. Muse and sol are nowhere near fable and Astra on the omniscience index