4 ms·
Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated? I'm also really skeptical of benchmarks that place any Haiku mo
by jchw 2mo ago
Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?
I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.
"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...
This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.
- rfgplk 2mo ago> Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated? The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.
- finaard 2mo agoInteresting - when Kimi K2.6 came out I switched over from Anthropic models, with at that time comparable to better results for me. I was using Anthropic via API, heavier months were roughly $400 worth of Anthropic tokens - I can get the same thing done via a $100 ollama subscription.
- sambusa_123 2mo agoJust read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/ https://reinvently.co.uk/tools/ed-o-meter/tests/ Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...
- nylonstrung 2mo agoI think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that
- zuzululu 2mo agoyeah that haiku really diminishes the claims behind the benchmarks. luna-max is significantly cheap and it is a strong performer but i dont see it on the benchmarks. I just get a feeling this site started with an intent to elevate Chinese models above the rest so wouldn't be surprised if the whole prompt sail was set to that tune