3 ms·
> beats Claude in our Cyber Benchmarks Beats which model in Claude? Whenever a "benchmark" doesn't put precise model numbers in their headlines I am immediate
by admax88qqq 3mo ago
> beats Claude in our Cyber Benchmarks
Beats which model in Claude? Whenever a "benchmark" doesn't put precise model numbers in their headlines I am immediately skeptical. Either they don't know the difference (bad) or they are benchmarking against weaker models (misleading, also bad).
It's like when studies say "AI is bad at X" and they used GPT-3.5 in current year.
- ls612 3mo agoOpus 4.8 according to TFA. Whether or not the safety guardrails were responsible for the difference is an open question but for a dev who wants to secure their software who doesn’t work at one of the blessed Glasswing companies it doesn’t really matter why, it matters what the best tool you actually have is.
- InsideOutSanta 3mo agoThey say "Claude Opus 4.8" in the first paragraph.
- crm9125 3mo agoWe're supposed to read the article? How are we supposed to stay skeptical of everything if we read anything!?
- simplyluke 3mo agoAnthropic's own models perform differently under the same version depending on how much they've decided to quietly downgrade them.