4 ms·
"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations" Then why does it have separate
by eckr 1mo ago
"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations"
Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??
- unglaublich 1mo agoMaybe they do that opaque degradation trick that whenever it's asked something questionable, it'll route to a worse model instead.
- manquer 1mo agoThe implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different. Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.
- Creamsicle47 1mo agoThe model cannot complete that task, for one reason or another, and therefore it scores lower.
- rcr-anti 1mo agoArtificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.
- iAMkenough 1mo agoMakes more sense if you recognize that Anthropic intentionally degrades outputs for most customers. Vetted customers get excluded from that practice.