5 ms·
The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M
by GodelNumbering 25d ago
The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M).
This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:
Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.
GDPval-AA v2: +1.5% vs Opus 5.
OSWorld 2.0: +2.5% vs Opus 5.
Humanity's Last Exam (with tools): +1.6%
Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?
- Tepix 25d agoDeepSeek V4 Pro cache read pricing is $0.022 (offpeak) and DeepSeek V4 Flash cache read pricing is $0.007 Makes it super affordable!
- nsingh2 25d agoFrom Artificial Analysis cost per task, it looks like Fable 5.1 (max) is more expensive per task than Fable 5 (max)? Cache hit price went down, but the other components still add up to more. Edit: 5.1-xhigh seems to be cheaper than 5-max, and 5.1-xhigh has a higher index score than 5-max. Also interesting that Fable 5.1 (high) is comparable to Opus 5 (max), but nearly half the price. https://artificialanalysis.ai/models#cost-tabs https://artificialanalysis.ai/models#cost-tabs
- GodelNumbering 25d agoInteresting, even if we were to ignore the cache-hits, reads and output, the reasoning cost (aka test time compute) per task should remain a fully comparable metric - it went from $1.25 (Fable5) to $1.48 (+18.4%) for an improvement significantly lower than 18%.
- nsingh2 25d agoI would expect the benchmark scores to be nonlinear near the top, as the easier tasks get solved and the harder ones are left over. So going from 10 to 15 would be easier than going from 60 to 65. I only take the Intelligence Index value roughly though. Considering they put Opus 5 (High) at the same level as Fable 5 (Max), I don't trust it that much.
- monkpit 24d ago> Considering they put Opus 5 (High) at the same level as Fable 5 (Max), I don't trust it that much. Have you used both? I’ve never experienced any seemingly greater level of intelligence from Fable 5 over Opus.
- glub 24d agoFrom my limited testing of just 2 hours, reasoning output of 5.1-max is at least 7x of 5-max, on the same project and comparable prompts. It reasoned for ~2 minutes trying to figure out an appropriate directory name. I've never seen 5-max do that. Could be a misconfiguration though.
- fny 24d agoI haven't dug deep but I burned through 30% of my weekly usage in a few hours which shocked me at first.
- oefrha 24d agoI did a bunch of Fable 5.1 xhigh review work on a bunch of critical components, ones with direct comparison from Fable 5 xhigh runs from two weeks ago. Token cost was 1.5-2x for each component.
- glub 24d agoYeah, after some more testing, I think I'm going to pin it back to 5. It feels like 5.1 is 5 that has higher reasoning threshold. I've been using fable as orchestrator anyway, so I see no reason to use 5.1.
- supern0va 25d ago>Has frontier progress finally stalled? It wouldn't surprise me if we start to see minimal performance gains from incremental changes to base models. It seems like the gains from the Opus 4.5+ incremental updates were a result of Anthropic learning a lot about post-training, the gains from RLVR, etc. If new post-training techniques are seeing diminishing returns, we could just be back to waiting for new large pretraining runs at larger sizes for gains (even if those ultimately end up getting distilled down into smaller models because the economics for serving anything larger than Fable isn't practical).
- anthonypasq 25d agoit seems to me that OpenAI is the only actual lab that truly understands reasoning. they have the best reasoning efficiency, they get pretty uniform improvements with more reasoning compared to other labs. (theres been plenty of graphs where models do worse with more reasoning), and i suspect their models are a lot smaller than we think. i think the next gen of openAI models are going to be quite insane tbh.
- rxyz 25d agoAnthropic did not get much bite because they don’t offer zero data retention with fable
- deleted 25d ago[deleted]
- z3dd 24d agoThis is exactly the reason why fable is blocked at the company I work at.
- 6thbit 25d agoHuh! Yeah that feels more like an opus5.1 than a fable5.1.
- arizen 25d agoPartially frontier moved to cost and speed axes.
- johnsmith1840 25d agoI'm a heavy user and fable is great the #1 reason I stopped using it was the horrible safegaurd filter. I found sol close enough in capability and have only been blocked when my request was an obvious offensive cyber work. Fable blocked me on almost everything. Optimizing a OS build? -> block Securing a container -> block 60% is nowhere near enough for that safegaurd system. This just means I am going to be blocked half as much? Any long running task will likely get blocked. Say you give a single big prompt and fable goes off for 6hrs of work. At hr 5 it gets blocked you now have the option of a much dumber model taking over and wrecking it or losing the entire 5hrs of work. That risk is beyond terrible and deffinetly not worth a 5-10% percieved improvement on my end. I previously would just bring sol in when that happened and realized sol is stupidly close in capability.
- andai 25d agoI never hit Anthropic's safety filter when I'm doing something illegal, only when I'm not.
- NewsaHackO 25d agoHonestly, unless what you are doing is frankly illegal, it usually is possible for you to get around most safeguards for coding things if you also know how to write code. Most problems have a separable completely innocuous core that Fable would gladly do. Then you can implement the problematic parts yourself. Particularly things like copyright issues, web scraping etc.
- glub 24d ago> Then you can implement the problematic parts yourself. Or with another LLM, but yeah. The only issue is when it's a monorepo and fable does ls/grep. I've got a file named `system_prompt` in a completely innocent project and as soon as fable accidentally stumbled upon it - cyber. Hacked together something with omp and sandbox-exec so that only whitelisted models can see some parts of the project. Works pretty well.
- nonethewiser 24d ago
- andai 25d ago> This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Does that mean that generally available intelligence is now constrained by Moore's law? We have to wait for the actual price to come down.
- andai 25d ago> Probably leaves no room to place Opus 5.1 anywhere. Well it'll probably be better than Fable again, lol
- black_knight 25d agoI haven’t yet had a week without spending my Max Fable allowance. For my work (formalised mathematics) Fable is my go to for hard(ish) tasks and problems – of which I have many! I hope they keep making it smarter! (Cheaper would be nice too, but smarter is my priority!)
- einsteinx2 25d ago> Has frontier progress finally stalled? From my experience using coding agents approximately 7 days per week for the past year and a half or so, we hit the top of the S curve about a year ago around Opus 4.5, and it’s mostly been harness and other tooling improvements since then with small percentage improvements coming from the actual models. I was saying this already months before Fable dropped and thought from all the Mythos hype that maybe I was wrong…then Fable came out and was barely better than Opus 4.8. Considering how many more parameters Fable is supposed to be than Opus, we seem to have hit a scaling limit at least with current transformer architecture considering how closely Fable and Opus benchmark and perform in practice.
- poink 24d agoAnecdotally, this has also been my experience Fable and Sol are better than Opus 4.5, but I don’t think I’d be weeks ahead on my projects if I’d had them in December
- anukin 25d agoThe issue with many of these benchmarks is that it doesn’t take into the real world usage of the model. Fable for me was a step above opus. The real reason I stopped using it is because of misanthropic. I was hospitalized and asked to extend my claim to fable credits and they responded to it by denying it. I regret buying annual plan instead of monthly one.
- unsupp0rted 24d agoIt's because GPT Sol is equally good and established a price ceiling
- jamiek88 24d agoSol is excellent. I’m amazed I can use it on my teeny $20 plan.
- m-schuetz 24d ago> Anthropic did not get much bite on Fable at its original pricing I stopped using Fable because it kept stopping itself due to safeguards.
- nonethewiser 24d agoDo you pay to get told no by Fable?
- astro1234 24d agoI don't think it's a stall, two ways I would believe there is a stall: - Does the epoch capability index progress show signs of plateauing? I consider this a good aggregate measure of diverse benchmarks into a single capability index. If we see things slowing down here thats a pretty direct and convincing piece of evidence for a stall. - Do we see any signs that scaling laws are beginning to fail? That would be by far the most alarming to me, since I would interpret that to mean that the entire premise of this unprecedented capital allocation tsunami is broken. Neither of these are true (for now). Progress is marching the same as it has for 4+ years now. It's still the same time to get a generation leap (I think like ~16-18 mo? Epoch has it) like GPT4->5. My theory is that people interpret plateauing because the releases are far more frequent now than they were in the past.