4 ms·
This whole thing immediately reads as Claude generated, making it hard to take seriously. Why do these results contradict existing serious attempts at benchmar
by hellohello2 1mo ago
This whole thing immediately reads as Claude generated, making it hard to take seriously.
Why do these results contradict existing serious attempts at benchmarking LLMs? Namely:
https://artificialanalysis.ai/ https://artificialanalysis.ai/
https://arena.ai/leaderboard/agent https://arena.ai/leaderboard/agent
- koe123 1mo agoHaha theres even an em-dash in the title
- ckocagil 1mo agoBecause it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.
- irishcoffee 1mo agoThe whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.
- hellohello2 1mo agoYes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.
- Tepix 1mo agoNote that in AAs report, Kimi K3 was at #1, then they updated their criteria and published a new report on the same day where it was no longer at the #1 spot. They may be under pressure not to declare a chinese model as #1.
- rdsubhas 1mo agoI'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style. This is the new witch hunt, or virtue trolling. Look at the later comments, they have substance oriented discussions. HN Mods - can you please consider a policy against such comments, it's overflowing the site and is diluting discourse and value. If people don't like an article, they can simply ignore it. These articles are reaching the top because enough people consider it of value. But when an article comes to the top, and the topmost comment and discussion thread is an unqualified witch hunt, it's getting sick.
- woodruffw 1mo ago[dead]
- Glyptodon 1mo agoI don't know about Claude specifically, but I think people are starting to internalize a sense of when prose reads as having "AI-smell" and I do agree phrases like "Same driver, same track. The LLM is the star." trigger it for me too. That said, that doesn't mean a ton about the whole thing - could be anything from humans starting to echo AI style to someone writing "give my results a headline summary" to an AI to someone saying "here's the data, write an article."
- rdsubhas 1mo agoAI is trained on human data. And high quality human data at it's best. Can we assume everything we think as AI - must have had a high-quality human pattern behind it, and there is no way to 100% prove which is which - unless the author shows a screencast of them typing the artice? This is not healthy. The right thing to do is – if someone doesn't like an article, they should ignore it – they shouldn't so confidently brand it AI without any proof at all, just because it fits their mood and style.
- adastra22 1mo agoAt this point AI is trained on AI data, and it is moving AI output into super attractors that have no correlation with human text.
- seizethecheese 1mo agoThese results don’t just contradict more serious benchmarks, they are wrong on an entirely different axis. This is a saturated benchmark. Haiku gets 96%. The results here are “not even wrong” and this being #1 on HN right now is a massive smell of either bots or massive ignorance or both.
- urams 1mo agoPeople REALLY want the open models to be better than the frontier labs'.
- geek_at 1mo agoinvestors REALLY want the closed models to stay frontier forevery. mmw the future of AI will be local and offline
- ed-is-ai 1mo agoRead the benchmark - it's on the basis of being 'good enough' for everyday tasks. Which is what fits most applications right. This is not a benchmark for testing them against an Einstein.
- SwellJoe 1mo agoWhile I was inclined to push back on the results, with Fable and Sol being so low, I have to admit I've also run into refusals several times since the latest models have arrived, and I've had to use Kimi K3 or DeepSeek to complete the task. Usually security auditing type stuff, but Fable balks at all sorts of ridiculous things, sometimes stupid things. I've even had Fable fall back to Opus and then Opus refused the task as well. So, it actually is becoming hard to use US models for everything because they refuse to work on a pretty broad selection of security and security-adjacent tasks. I guess if you're not at a Fortune 500 or a member of a fascist government, you don't get to use the best models to protect yourself and that's just how it's going to be. But, you're right. The prose is miserable Claude-speak, difficult to wade through.
- hellohello2 1mo agoInteresting idea, I had not considered refusals. I have ran into some as well although rarely. I'm not certain the 28 tasks described would trigger it though, if I understand correctly the security tasks are about avoiding prompt injection and not about doing security work. EDIT: You were correct, Fable and Opus reject some of the coding tasks, which is why they score lower. Thanks for explaining. EDIT2: I believe this benchmark is invalid, my Opus 5 runs the supposedly rejected tasks just fine.
- EddieLomax 1mo agoIt's because it was written by Claude Fable 5. https://reinvently.co.uk/about/ https://reinvently.co.uk/about/ > Fable Anthropic * Research Editor * Claude Fable 5 > Challenges claims, tightens methods and prose, applies British English and removes hype that the evidence cannot support.