4 ms·
>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. It's just inadequate
by logicchains 2mo ago
>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression.
It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
- Jensson 2mo agoYes, and good senior software engineer is ahead of fable, but benchmarks can't capture that either. We already know from testing humans that test scores don't correlate that well with how effective a person is at work. Same applies here, we just aren't that great at making good tests.
- otterdude 2mo agoIf most models were getting 100% on the test it would be an inadequate benchmarks. What were seeing is all models failing to ace these tests. "Benchmark Saturation" is term that promotes lowering the bar.
- deleted 2mo ago[deleted]
- Planktonne 2mo ago> Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this. I have seen people claim the exact opposite. If there was so much progress, then you wouldn't have endless disagreements with people championing their own favourite model as the strongest.
- ofjcihen 2mo agoThe two times I tried to use fable I had it attempt something I had already had Opus 4.6 do with no issues. It blasted through 10% of my weekly allowance on a 100$ a month sub and produced something broken and nonsensical.