3 ms·
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is tha
by 6thbit 2mo ago
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
- gallerdude 2mo agoCapable in term of AI R&D, not capable in terms of hacking (which caused all the Fable drama.) But agree, confusing wording.
- square_usual 2mo agoEasy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.
- llelouch 2mo agoYep , same with 5.6. Fable is still the best.
- lifty 2mo agoBut still nerfed compared to the initial release.
- solenoid0937 2mo agoOnly when you hit the cyber classifiers.
- 6thbit 2mo agoHonestly that's the simplest explanation and thus likely the correct one.
- usef- 2mo agoThey don't have a track record of benchmaxxing. The simplest explanation is that the blog post lists the things this model is better than Fable at, but not the things it isn't.
- eli 2mo agoThat's a plausible explanation but I'm not seeing evidence for it. I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
- matt2000 2mo agoThis is an interesting idea, without giving away your benchmarks specifically what kind of stuff do you test? I might try to assemble something myself, it's so hard to determine model quality from the system card these days.
- eli 2mo agoI have actually been planning to open source the framework. Only benchmarks I care about are the ones that look like real work I do. So it makes it easy to trawl your own repos looking for benchmark task candidates from real bug fixes or features. Then as a bonus I can compare the agent’s solution to my own as a reference. So most look like that but I did include a few one-shot “build an app that solves this problem” and some qualitative design tasks and a tough algorithmic optimization one.
- matt2000 2mo agoWould open sourcing it make you feel like the results might be compromised? I'd rather keep it private so it's specific to my use cases and not included in any training date (no matter now small the signal in the overall data).
- eli 2mo agoOh sorry I meant the framework. Though I could see posting a couple of the tasks.
- bisonbear 2mo agoAlso working on a product to build tasks from your own work for testing coding agents. Main thing I would offer is to look carefully at the agent trajectories - they love to figure out ways to cheat. Additionally, consider what "winning" means. If just using test pass rate, consider that tests might not encode what good means in your repo. I have been having success using "equivalence with merged PR" as judged by an LLM as a signal.
- flakiness 2mo agoMaybe they don't want to say that to avoid the government scrutiny.
- HarHarVeryFunny 2mo agoIt seems they are trying to thread a needle here - they want to say it's very strong, but apparently this time do not want to invite extra government scrutiny. They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard. "As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."