3 ms·
So it's a Fable class LLM? DSV4Pro vs Fable5 HLE w tools 60.0 vs 63.0 Terminal Bench 2.1 87.9 vs 88.0
by bel8 2mo ago
So it's a Fable class LLM?
DSV4Pro vs Fable5
HLE w tools 60.0 vs 63.0
Terminal Bench 2.1 87.9 vs 88.0
Cybergym 83.3 vs 83.1
DeepSWE 62.7 vs 70.0
Toolathlon-Verified 74.1 vs 77.9
AutomationBench (Public) 31.8 vs 29.1
DSBench-FullStack 71.1 vs 77.2
DSBench-Hard 67.2 vs 68.3
- aftbit 2mo agoFabble lol
- qiran87 2mo ago[dead]
- eli 2mo agoFable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5
- wren6991 2mo agoWe have a first-party figure from the system card [1]: > Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash). So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5. [1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e908266086df425cdf5/Claude%20Fable%205%20&%20Claude%20Mythos%205%20System%20Card.pdf https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...
- nikcub 2mo agothat DeepSWE result is likely most indicative of how you'll find real world usage