21 ms·
Claude Fable 5.1 and Claude Mythos 5.1
What's new in Claude Fable 5.1
– https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1 https://platform.claude.com/docs/en/models/fable-5-1/whats-n...
System Card: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32...
- genxy 1mo agoThis change was for them, not us. I am touching grass until next week while they get this shit sorted out. Not on my time.
- Yash16 1mo ago[dead]
- zb3 1mo agoSo bullshit safeguards are still there.
- ghoshbishakh 1mo ago"Content provenance" seems to be activated with this model.
- nezhar 1mo agoThis time it came with a usage reset
- sva_ 1mo agoGreat, my usage reset is in 10 hours ... And my 5 hour window was due to be reset in 2 hours (barely used), now its in 5 hours - so this reset effectively gives me 1 less 5 hour reset for this weekly cycle.
- rirze 1mo agoSame... I think this timing aligns with a reset they gave months ago.
- jonesy827 1mo agoUnless you run overnight, you could schedule a cron job to send a basic claude -p prompt such as "reply with hello" using haiku to align your usage windows. That's what I do.
- sva_ 1mo agoYes something like that is what I did, so I had my 5 hour reset window to be at 2 hours so I could work. But anthropic reset it so it went back to 5 hours.
- jorl17 1mo agoThis was the best thing for me. 98% Fable usage resetting only Thursday and just got this early. Couldn't be happier.
- scronkfinkle 1mo agoHas anyone been able to get anything substantial done with Fable in the first place? I more or less had totally given up on using it since the alignment checks were so sensitive that it pretty much always threw me back to Opus.
- vessenes 1mo agoCombo of that, laziness and load-bearing language + the penchant for making up weird dense conceptual names pushed me to sol 5.6. They seem to indicate it is a less annoying writer in the announcement so I’m curious to try it out, though. Ironically one of their demos is speeding up inference - do us normies get to do that with Anthropic tech??
- UltraSane 1mo agoI have used it to write some scripts but it is incredibly expensive.
- prettyblocks 1mo agoIt's probably a good model for folks doing basic software stuff, or humanities related tasks, but I work in cybersecurity on the defense/detections side and I haven't been able to use it for anything even with being in the CVP. It downgrades to Opus every time.
- ronsor 1mo agoFable is way too expensive for basic software stuff. Other models are more than good enough for that. In general, if Fable isn't blocking you, there's a high chance a lower tier model would work fine.
- elevation 1mo agoThe only problems Opus struggles with, Fable won't take on. I was porting some software from Win32 to linux. Opus was running in circles. Fable was going great until it saw some authentication code and bailed.
- 1mo ago
- simonw 1mo agoBit of a discount if you're using caching: > same input and output prices, with cache reads at a quarter of the cost This should impact any long-running agent since subsequent calls can benefit from cached reads for previous transcripts.
- Twixes 1mo ago~30% reduction in real-world task cost vs. Fable 5 in our evals at viktor.com ! Caching goes a looong way
- behnamoh 1mo agoAnd yet, despite this, the quota limits went down by 17%.
- davely 1mo agoIn my opinion, this is a bit disingenuous. They were _temporarily_ increased in May by 50% [1]. They continued to extend them through July and August (admittedly, their messaging around this has just been a complete mess and they frequently pushed the deadline back as it approached). So, now they are giving you a 25% quota increase compared to where things originally stood in May. So, let me ask you this: assuming you knew that the 50% quota increase was temporary all along, would you then have complained about Anthropic restoring things back to the original limit? [1] https://www.anthropic.com/news/higher-limits-spacex https://www.anthropic.com/news/higher-limits-spacex
- Petersipoi 1mo agoOn the contrary, you and Anthropic are being disingenuous by pretending that a usage reduction is actually an increase. Especially when the 20x max plan isn't actually anywhere near 20x, as people have recently realized.
- anthonyrstevens 1mo agoYes, some people will complain about anything (and everything) related to AI. And relentlessly push the most negative interpretation of any datum.
- spicypixel 1mo agoYeah but haiku 5 when?
- eshack94 1mo agoAsking the real questions. I've been wondering what the holdup on that is. Does anyone reading this have additional knowledge or insight on this?
- Jcampuzano2 1mo agoThey haven't really mentioned practically anything about Haiku in quite a while so I imagine nobody except for people inside Anthropic will have any indication. Maybe it'll come out eventually but they don't even include it on some of their comparison benchmarks anymore, so I figure its very low priority for them.
- wahnfrieden 1mo agoMore likely that they are embarrassed by how their attempts compare with OpenAI's Haiku analog, Luna.
- ayewo 1mo agoNot so sure since Anthropic has 4 model families while OpenAI has 3 for GPT-5.6. Claude Fable/Mythos vs GPT-5.6 Sol Claude Opus vs GPT-5.6 Terra Claude Sonnet vs GPT-5.6 Luna Claude Haiku vs ?
- HDBaseT 1mo agoThe pricing on Luna is just insanely good. Haiku simply cannot compete at almost any intelligence level against DeepSeek V4 Flash or Luna.
- wahnfrieden 1mo agoNo, you're off by one - likely because the OpenAI models punch above their weight in those comparison, hence my original message. You've shifted the comparisons to favor Anthropic. Fable/Mythos are much larger than Sol. They match to Astra which is supposedly at least 10T. Astra is already publicly confirmed as a new family. Opus matches to Sol. Sonnet to Terra. Haiku to Luna. Anthropic is able to compete at the frontier high-end by launching massively large expensive models. But their inability to compete on small models belies their efficiency aspirations across the stack.
- maxdo 1mo agoTbh with that price , not even willing to try . What are the benefits for a regular coding agent ? I barely have any errors already with 4.8 level , eg grok 4.6 , gpt 5.6 sol/terra behind router . Why do I need to pay so much money for this ? Any reason ?
- Philpax 1mo agoMaybe you don't! It is very possible that your problems don't actually need frontier-level artificial intelligence.
- deleted 1mo ago[deleted]
- tripleee 1mo agoI can't tell if this is insulting or not lol
- Philpax 1mo agoInterpretation is in the eye of the beholder :-)
- maxdo 1mo agoI do agree , it could be an insult to any software project probably ? But I do value more speed of iteration/verification cycle vs another 3% in cursorBench . At this point it’s business logic not the code that caused me troubles and extra thinking
- enraged_camel 1mo agoFable is significantly better at helping me think through (and untangle) business logic problems as well. I actually rarely use it for implementation because Opus 5 is good enough for my use cases. YMMV.
- taberiand 1mo ago
- re-thc 1mo agoThe biggest change is the price cut of course.
- pookieinc 1mo ago"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%." Glad to see this!
- benjiro29 1mo agoI hope that this also applies to the Subscription usage. As that can then stretch out Fable usage by a lot more.
- Narretz 1mo agoSounds like it specifically does not apply to subscription usage.
- re-thc 1mo agoThat's load bearing!
- LtdJorge 1mo agoBut does it fail open or close?
- fearmerchant 1mo agoThe gate is green
- aprilnya 1mo agoMy understanding is subscription usage generally has free cache reads, but I'm not sure if maybe Fable was different in that regard.
- MitziMoto 1mo ago
- robinpie 1mo agodo you think we'll go a full year without a new haiku lol
- porridgeraisin 1mo agoWith compute crunches and everything I am not sure it makes sense for anthropic to commit to haiku as an endpoint and thus a product. There is no telling they aren't using a similarly sized model behind their existing opus/fable endpoints for various subagent / summary purposes of course.
- 2001zhaozhao 1mo agoThey really should launch a new Haiku to compete with Luna imho. Luna is insanely good for the cost and it's my go-to for high volume batch tasks now.
- porridgeraisin 1mo agoTrue, but I am not sure how much uptake it has in their enterprise accounts. Slowly they are all coming to only care about that.
- cromka 1mo ago"with cache reads at a quarter of the cost" OK, I think that's what they meant when they suggested reduced extra promo usage will not sting this much.
- eis 1mo agoI am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big differences? I did notice Opus maybe making more mistakes repeatedly but I don't have hard numbers on this. I hope Fable 5.1 brings noticeable improvements. I am giving it a go now on my 20x Max plan on a problem that Opus 5 has struggled for more than week now and has made very slow progress with regular regressions on the way.
- rfgplk 1mo agoOpus 5 is better than Fable 5 except for creative programming work (like graphics). Fable 5 might be slightly better but the token cost isn't worth it.
- bredren 1mo agoHow are you evaluating the models? On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal. I primarily use the models via interactive sessions enhanced with custom tools and skill. For that Opus 5's benchmark superiority has not materialized into greater productivity and frankly has been quite a let down. The outputs are too often unreadable even after adding recommended prompts. There is an ongoing problem with the heron_brook system prompt affecting orchestration. [1] I've used Opus 4.8 since the second week Opus 5 was released. Over this time, Fable 5 has been reliably fantastic. Both in planning and direct execution on complex changes across code and infra. I'm a bit surprised that there doesn't (seem) to be a section discussing ~performance across different modalities. This system card and blog post too-often default to an API-based use case when the gander primarily experience Anthropic's models via interactive sessions. I understand waiting to comment until Opus 5.1 is available and handles these problems, though I am hopeful that Anthropic will confront the elephant in the room on Opus 5's failure to delivery great interactive sessions and the widespread negative feedback on the release. It would show the org is paying attention, taking steps to balance model evals between interactive and API use. Also, some empathy for customers that wasted time trying to make opus 5 work for them. [1] https://github.com/anthropics/claude-code/issues/80988 https://github.com/anthropics/claude-code/issues/80988
- canadiantim 1mo agoThank the heavens for quota resets
- niteshpant 1mo agoI don't know how I feel when all the documentations are written by AI for humans. AI to AI doc share: sure, do what you please. AI to human: please make it legible and flowly. example, "Every thinking block records which model produced it, and it's preserved in one direction only: Claude Fable 5.1 reads earlier models' thinking blocks, and no earlier model reads Claude Fable 5.1's." is a very Claude-isk way of writing. Choppy, long, and lacking flow.
- Bluestein 1mo agoUnless these people start offering free, unlimited inference for a cautionary period so we can test the new model without an up-front (re-)investment, I am not touching this load-bearing pile of neuralese spew with a ten thousand token pole.-
- anthonyrstevens 1mo agoAre you going to post the same comment on every AI article on HN? How is this useful to the discussion?
- Bluestein 1mo agoThis, is the first time I have not only made this comment, but opined on this issue at all, actually.-
- Bluestein 1mo ago(The funny thing is that the comment itself was highly resonating, until the PR - or Claude! - came in, after a lag).-
- aennassiri 1mo agoYou tend to take a rather aggressive approach when commenting in here.
- sunaookami 1mo agoSadly still not available for Pro subscription. At least they reset everyone's limits.
- rirze 1mo agoThat feels bad, my weekly limit was going to reset today. (I wonder if mostly everyone's reset day is today as well...)
- sunaookami 1mo agoMine reset yesterday but I won't complain since I profited from the last two resets that were on Friday :D
- stillpointlab 1mo agoI didn't see it anywhere on their announcements, but when I restarted Claude (on a Claude Max account) I see the model is now Fable 5.1
- tarr11 1mo ago“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that humans have a token limit too and it can be exhausting to read Claude’s output. Prose density is not the same thing as succinctness.
- hungryhobbit 1mo agoAmen. I would trade some stupidity (say ten points on any benchmark) in exchange for a version of Opus or a similar model that actually gave me direct, concise answers.
- atishaykumar 1mo agoYou can possibly give instructions on how to respond to your questions.
- hungryhobbit 1mo agoYes, and they will work ... for like two turns, after which Claude will go back to its usual wall of text. And yes you could add context (memories, rules, CLAUDE.md entries, etc.): they won't help (for long). Same for hooks that remind Claude to be concise: it gets "attenuated" and starts ignoring any such instructions quickly. There's also writing guidelines ... but they're basically just more context with slightly higher weights (ie. Claude will still ignore them). I've even gone so far as to make a hook that identifies long responses and requests shorter versions (which is challenging in itself, as you need to run another lower-powered model to evaluate how long is "too long", as what's "long" when the expected answer is one line is different from what's expected for a ten line answer). However, that just shows you the long version, then some hook text, then (10-15 seconds later) it shows the short version. So I created a proxy that hid the long version/hook text for me ... but I had to abandon it because all that used up so much usage I was running out. I'm fuzzy on the details, but Caveman somehow "hacks" Claude in a way that gets past all that ... but it takes things too far in that direction, with "cave man" speech that sucks.
- exabrial 1mo agoDid we get thought traces back? If no, it's useless.
- BoorishBears 1mo agoLol we got literally the opposite: > *Fewer progress updates during long tool runs.* > The model writes less user-facing text between tool calls, especially at higher effort. Set thinking.display to "updates" (beta) to receive the progress updates it does write, and remove any prompt line that tells it to hold findings for the final response.
- sevenseacat 1mo agoI love when I make a request or ask a question, and Claude Code immediately queues up a dozen tool calls to edit files instead of explaining what it wants to do or why, despite repeated constant reminding that I will not approve tool calls without context and reasoning
- vandopereira 1mo agoIve been through this a lot myself, and thats why im building this tool called kaplira, it is an control layer for ai coding agents. it does not allow ai agents to touch files outside a determined scope, it learns as u code, learns from past mistakes, regressions, it injects cirurgical memories into the ai agents. u can download it for free at kaplira.com
- BoorishBears 1mo agoAt least half the changes are just anti-distillation strategies...
- zb3 1mo agoGood to know they're getting desperate, the sooner they implode the better
- 2001zhaozhao 1mo agoI really don't think they can stop it, only make it somewhat more expensive. As long as the model need to make tool calls on the user's computer, the user can record the trajectory and use it to reinforce another model to follow the same trajectory.
- as12fj 1mo ago"Democratization" through AI means that everyone has to pay a monthly Anthropic tax and only a small secret guild gets access to the real model. Jane Street is a partner? How sad indeed. Anthropic could front run them because they leak all the data.
- mlaux 1mo agoLooks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from other models verbatim (since it can see the decrypted version). I get that in their eyes it's an "exploit" but still kinda disappointing that they patched this
- giancarlostoro 1mo agoTo be fair, I assume they want to hide that not from their customers, but adversaries who use the way Claude models think and reason to refine their own models.
- appplication 1mo agoI have a hard time believing whatever prompts get Claude to reason can stay relevant secret sauce for long anyways. It’s not hard to A/B test something that gets you close enough, and it’s not Ike anthropic has uncovered the global optima of reasoning prompts.
- verdverm 1mo agoI don't really want the models I use learning from Claude at this point. Open weight models of similar scale are available now too, so I expect this "distillation"/"stealing" chatter to wind down.
- hyperpape 1mo agoMaybe, but that's sort of begging the question that those open weight models aren't significantly trained using "distillation"[0] [0] not technically distillation. https://thomasdullien.github.io/posts/2026-06-15-rl-economics-morally-charged-terms-and-distillation/ https://thomasdullien.github.io/posts/2026-06-15-rl-economic...
- sashank_1509 1mo agoReads like AI slop, surprised they can’t see it in their blog post. No human wants to read in such prose
- dainiusse 1mo agoDon't care unless it is priced in as other models.
- tclancy 1mo ago> Denser prose in places. Really? Interesting choice. Pretty much every CLAUDE.md file I have starts with something about Hemingway, terseness and treating every word you use like you're carving it on your own back, but different strokes for different folks. I suppose I haven't heard from anyone who enjoys how wordy Claude is because they aren't done writing their post yet.
- ramon156 1mo agoI yearn for a model that can churn through claude text and write sensible text. so far gemini is pretty good at that, even in the low variant
- laacz 1mo ago"Explain in simple terms" works.
- MadsRC 1mo agoI was looking forward to using Fable for cybersecurity work, but kept getting bumped to Opus… Signed my org up for CVP, went through the trouble of procuring a separate team plan from our main org as Anthropic can only disable cyber safeguards for an entire org and not individual users… After months of trouble dealing with KYC and procurement I finally got CVP for my security org and today I found out that CVP (which is what removes cyber safeguards) does not apply to Fable… So yeah, unless you’re a Project Glasswing member, there’s no using Fable (which with Glasswing is Mythos) for security work… Absolutely useless… Didn’t they just sign some “we must use AI for cyber defense before the bad guys do” and then they artificially cap us by not allowing Cyber-unlocked Fable… Sigh…
- hungryhobbit 1mo agoSome of this is Anthropic, and some is the Trump administration ... ... but some is definitely Anthropic, so I'm not trying to let them off the hook; I'm just pointing out that the government is partly responsible.
- seaurchinzee 1mo ago"Cache reads now cost 75% less, or $0.25 per million tokens." For me, at a typical 95% cache hit rate, I think my optimal context window size before autocompaction goes from ~200K to ~400K tokens. Great for longer horizon tasks.
- cute_boi 1mo agolooks like it is only for api.....
- kingstnap 1mo agoDo API prices not affect usage limits for subscriptions? They do in Codex.
- seaurchinzee 1mo agoOh dang, that's really unfortunate, nice catch. At least Claude subscription users got a usage reset. But yeah, I can't help but feel Codex is far more generous with their subscription quota at the moment. I've been using Fable to orchestrate GPT Sol Max and Sol Ultra agents all day, and I've barely made a dent.
- eaf7e281 1mo agomay i ask where did you get this? i try to look through the docs, but i didn't find where they said its only for API is it in the system card? really hope not, that change the only positive part in this release
- demibabs 1mo agoThey specifically said it in the press release. I don’t see why they wouldn’t have mentioned it if it also applied to subs
- skiph 1mo ago[flagged]
- rybosworld 1mo agoInstead of a new model that's going to have unreasonably shallow usage limits, I wish they would: 1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan 2) either fix opus 5, make it completely free, or delete it entirely
- Someone1234 1mo agoI'm legitimately out of the loop; what is going on/broken with Opus 5?
- sidrag22 1mo agoit just doesn't interact good with human beings, and it leaves incredibly strange long winded comments within code filled with session context that will likely not be relevant later on. Also always seems to have this annoying tendency to leave "questions for you" at the bottom of every output. Just a high friction human interaction type model, imo should never have even been released, regardless if it scores better on whatever tests, its a horrible experience and a downgrade over past models.
- tstrimple 1mo agoI have to wonder if everyone else is just running these models raw without any custom instructions. I hear all these things about voice and code comments and those are all things I've dealt with long ago via claude.md instructions, rules, and hooks. My claude can already respond in any "voice" I want and the quantity and quality of comments is within my control.
- sidrag22 1mo agoit really just seems like people pump out that its on the end-user, and i just disagree. They have a walled garden around claude code and using their models within it, it should work instantly out of the box when going from an opus 4.8 to an opus 5.0 with the same workflows. it doesn't. claude.md for all my projects are fairly tight, its seldom where im upset at anything a model does, and if it happens, its likely because i swapped provider and didn't realize i was failing to feed it proper context beforehand. Opus 5.0 fails in different ways that I haven't had to deal with. Its insufferable with its choice of language, something I've never had to compensate for on any other model across any provider, so of course I have no preexisting rules for that, it also is sometimes just incredibly stubborn and just WONT finish, and requires several just "keep going" prompts. This is much different than the issues people would make fun of users for in regards to treating models like slot machines and just pulling the lever over and over, this is more its stopping for no reason short of its task, and literally just needs to be told to continue? absurd. Most of my workflows have reference material, with standards set, why opus 5.0 is the only model that fails to follow those standards and inserts wildly long weird code comments is not a failure on the end-user, thats the model failing. I can be MORE explicit of course, but i shouldnt need to be, this is supposed to be 5.0, its a downgrade. I went back to 4.8 and all these issues vanished.
- 2001zhaozhao 1mo agoThere's now a 40X discount in the cache input pricing instead of 10X. This seems to point to them having achieved some kind of optimization in attention mechanism perhaps along the lines of DeepSeek V4, which had a similarly high discount between cache input and normal input. In real world use, the savings should be quite noticeable. For example, you can now use the model at 800K tokens context window at the same cost efficiency as the previous model at 200K tokens context window.
- felixrieseberg 1mo ago(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful. [1] https://github.com/harbor-framework/terminal-bench-science https://github.com/harbor-framework/terminal-bench-science
- behnamoh 1mo agoAt this point, I don't believe a word from Anthropic employees; you guys have lost all the goodwill that you accumulated over months last year.
- chews 1mo agoI share this sentiment, I really did like the models... then the finger printing, encryption of thought traces, staggered access, the constant NO's from Fable on cyber related issues for looking at bugs in my own code... I'm glad I swapped to Kimi/GLM... now with the deepseek harness, I don't even miss Claude Code. I really hope open models give them the market reckoning they wholeheartedly deserve.
- nullstyle 1mo agoHave you shared any details about your dsh setup anywhere? I’ve only dipped my toes in and would love someone else’s perspective on how they use it
- chews 1mo agoI've not, but really should. I run it on exe.dev, it's an ephemeral VM company and they have an agent of their own called shelley (which I used locally as well), Having kicked the tires on DSH(deepseek harness), I ported Shelley's skills into DSH, they are pretty simple text files that were easy to bridge over, it is more verbose but the plugin nature of it was really easy to extend, for example, I built a plugin that checks my claude usage windows and when I get to 80% stop asking new agents for help.
- jumploops 1mo ago> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.
- vablings 1mo agoSoftware will be buggier than ever but also way less buggy.
- unglaublich 1mo agoIt's going to be 50% less buggy, but we're going to write 10x as much code too.
- alasano 1mo agoGood software will be good-er. Bad software will be nightmare fuel.
- vablings 1mo agoI think bad software has the possibility of redemption with rewrites and re-engineering efforts. For those of us who are license locked that's probably never going to benefit us :(
- farkerhaiku 1mo agobad software will be replacable.
- pphysch 1mo agoBad software, as in stateless programs, doesn't actually matter and never did. They can be replaced trivially. The problem is the real world isn't made of stateless programs, but lots of important data in bespoke formats/schemas, and if you change the shitty software that interacts with the important data, in the wrong way, you can lose everything.
- 2001zhaozhao 1mo agoHi Claude, please cure aging, make no mistakes
- rvz 1mo agoEver since this "comedic incident" [0] you are apparently "not allowed" to make this specific joke as you are going to "upset" some people who don't get it. /s But eventually AI will cure something, unironically. It may be Claude, or another AI company. [0] https://news.ycombinator.com/item?id=48838228 https://news.ycombinator.com/item?id=48838228
- as128ah 1mo agoI'm sorry, this feature is only available to project Glasswing members for safety reasons. Would you like a port of Emacs to Visual Basic instead?
- mapontosevenths 1mo ago> Hi Claude, please cure aging, make no mistakes Done. The average human lifespan is now zero.
- deleted 1mo ago[deleted]
- ckugblenu 1mo agoThis coupled with verification primitives will be quite compelling. we really have to start reimagining existing systems and processes from the ground up.
- dfltr 1mo agoThis feels kind of petty, but what is going on with those fuckass clouds in the background? Did no one notice how uncanny that whole thing looks?
- tosh 1mo ago> The watermark doesn't change the meaning, quality, or readability of the output how?
- apsec112 1mo agoHere's the paper describing the technique: https://www.nature.com/articles/s41586-024-08025-4 https://www.nature.com/articles/s41586-024-08025-4
- deleted 1mo ago[deleted]
- jampekka 1mo agoIt manipulates the PRNG seed in a systematic way, keeping the same token sampling distribution. https://www.nature.com/articles/s41586-024-08025-4 https://www.nature.com/articles/s41586-024-08025-4
- philipwhiuk 1mo ago> Claude Fable 5.1 follows explicit tool instructions reliably. Moving stuff out the API into prompt engineering is obviously less reliable but necessary for progression to 'actual intelligence'. Will be interesting to see if it really is solid.
- deleted 1mo ago[deleted]
- lousken 1mo agoWhile haiku is almost one year old. What a joke
- mentalgear 1mo ago> Forced tool use is not supported That seems unfortunate for 3rd party integrations that expect stable output - what that really necessary ?
- thisisauserid 1mo agoZero data retention coming soon! ... with the condition that you store 100% of your data and make it available to the US government and possible others.
- _islo 1mo agoI’m really excited to try this out. Fable and Opus 5 constantly wow me when working together. Unfortunately, I’m a little burned because of technical issues. Anthropic accidentally over-billed my account, and when I reached out to the support bot, it downgraded my account to a Free account. It’s been impossible to get it resolved and I have almost $200 held hostage. I don’t want to do a charge back. I’m one of the main advocates for Claude Code at work, I use this subscription to try out new features before it’s available at work. The whole experience has been illuminating about our dependencies on these AI companies.
- maxgee 1mo agohad a similar issue. just do a chargeback.
- mannanj 1mo agoyou aren't the only one with this issue. many other people I've heard had a similar issue with anthropic billing. I also had a weird edge case behavior around billing where it blocked my usage due to an unpaid bill but then also wanted me to pay for that blocked unavailable usage when I would reinstate my account. I am disappointed in how anthropic handles billing, and is using AI sloppily for customer service around here. Very unprofessional, and at this point since its been well known and shared, it also is feeling unethical.
- mannanj 1mo agoand after posting this, and recently restarting my anthropic subscription: - I see 2 billing placeholders on my bank account - 1 subscription actually went through - my bank flagged anthropic's subscription initially as suspicious and I needed to verify it Is this normal for anyone else? hasn't happened with my codex subscription.
- ayhanfuat 1mo agoI noticed they reset the usage and I was kind of happy because this week it was using my quota much faster; I assumed they fixed that. Apparently it is for the celebration of 5.1?
- hungryhobbit 1mo agoThey dropped your usage limit by 17% this week .. They claimed to "raise" it, because they did ... while also removing the temporary increase they applied for a few weeks ... but the net effect is you can use 17% less than you could last week. On top of that, recent versions of Claude had a ton of tools added, and all those tools use up significantly more context/usage than before, so the moment you open a Claude session you are already using a lot more (I forget how much more) usage ... just to do the same exact thing you did last week.
- joshfraser 1mo agothe counterbalance to the AI doomers has always been the fact that everyone has equal access to AI. i hate this new world where Anthropic believe they should be the ones to decide who gets access to super intelligence and who doesn't.
- skiing_crawling 1mo agoAll the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.
- jdgoesmarching 1mo agoAll the benchmarks in the world don’t matter if the subscription forces you into a walled garden of slopcoded apps. I’ll stick with Codex and, increasingly, open source SOTA models.
- enraged_camel 1mo agoGreat, thanks for sharing.
- purpleidea 1mo agoI notably had an issue that it wouldn't work on a "remote execution" (running a command over SSH) coding problem until I did a sed to remove the word "execution". Incredibly dumb. I'm not doing any murders. Easiest to just switch to the Chinese models.
- arizen 1mo agoThe only company to use Claude.md instead of Agents.md standard
- mirekrusin 1mo agoWith new watermarking you may now get Hullaballooing.md
- celrod 1mo agoI'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.
- tstrimple 1mo ago
- apt-apt-apt-apt 1mo agoI'm so suspicious of this after Opus 5 benchmarks scored it higher than Fable 5, yet Opus 5 was untrustworthy (overconfident, error-prone).
- GodelNumbering 1mo agoThe price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement: Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5. GDPval-AA v2: +1.5% vs Opus 5. OSWorld 2.0: +2.5% vs Opus 5. Humanity's Last Exam (with tools): +1.6% Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?
- Tepix 1mo agoDeepSeek V4 Pro cache read pricing is $0.022 (offpeak) and DeepSeek V4 Flash cache read pricing is $0.007 Makes it super affordable!
- nsingh2 1mo agoFrom Artificial Analysis cost per task, it looks like Fable 5.1 (max) is more expensive per task than Fable 5 (max)? Cache hit price went down, but the other components still add up to more. Edit: 5.1-xhigh seems to be cheaper than 5-max, and 5.1-xhigh has a higher index score than 5-max. Also interesting that Fable 5.1 (high) is comparable to Opus 5 (max), but nearly half the price. https://artificialanalysis.ai/models#cost-tabs https://artificialanalysis.ai/models#cost-tabs
- GodelNumbering 1mo agoInteresting, even if we were to ignore the cache-hits, reads and output, the reasoning cost (aka test time compute) per task should remain a fully comparable metric - it went from $1.25 (Fable5) to $1.48 (+18.4%) for an improvement significantly lower than 18%.
- 1mo ago
- koolba 1mo ago> Data retention. Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention. This is interesting. I wonder if customers will be allowed to create an auto expiry for their own data to prevent future subpoenas. That’d be a treasure trove for discovery.
- stillpointlab 1mo agoMy only concern is that sooner or later the best models will be priced out of my ability to pay. I have been happy with Fable 5, it has done great work for me so far. Very excited to try out Fable 5.1 and see what differences and improvements there are.
- bhelkey 1mo ago>Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token.
- eckr 1mo ago"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations" Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??
- unglaublich 1mo agoMaybe they do that opaque degradation trick that whenever it's asked something questionable, it'll route to a worse model instead.
- manquer 1mo agoThe implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different. Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.
- Creamsicle47 1mo agoThe model cannot complete that task, for one reason or another, and therefore it scores lower.
- rcr-anti 1mo agoArtificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.
- iAMkenough 1mo agoMakes more sense if you recognize that Anthropic intentionally degrades outputs for most customers. Vetted customers get excluded from that practice.
- simonw 1mo agoPelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F95ccf9b75804a7a7e1d7d9e106a89caa https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Facf6ab2516527d97f04b9f07d61a7cad https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30! Excerpts from the reasoning trace: > Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter. > Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space. > I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...] > I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...] > Now I'm checking the vent line placements on the helmet, making sure they sit far enough inside the helmet's edge given the stroke width and rounded caps, and confirming each vent stays within the helmet's circular boundary. [...] > I decide skipping a handlebar bell and tire highlights since they're unnecessary additions. Now I'm reconsidering the front fork's curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork's lean. This is a notable result because most of the recent Claude models have been pretty bad at drawing pelicans, at least when compared to models in the Gemini or GLM series.
- enraged_camel 1mo agoIn a way, this is the only benchmark I care about now. :)
- dabinat 1mo ago> This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their organization, or their conversations with Claude. How does this work if it doesn’t change the output?
- unglaublich 1mo agoIt does change the output, they never said it did not. They said it would not _noticeably_ affect performance.
- sroussey 1mo agoWhich is a lie. Or vacuous statement as Claude might say these days.
- anthonyrstevens 1mo agoHow is it a lie?
- sroussey 1mo agoThe performance of the output is poor and I can detect the watermark but only in a domain i am in the middle of, if it writes a summary of an email i can not tell, but in claude code in my codebase i definitely can tell. I am complaining about it!
- sroussey 1mo agoOr when it tries a different language to explain a word: "The four items the ledger had标 standing open"!
- Smaug123 1mo ago
- abroszka33 1mo agoLooks like agentic coding plateaued, and agentic scientific research is the new hype?
- mohitpaddhariya 1mo agoInterestingly, Claude’s output is now actually readable with Fable 5.1. Pretty sick.
- TuxSH 1mo agoUnfortunately isn't included in subscriptions and requires usage credits...
- enraged_camel 1mo agoInteresting that they seem to have gone all-in on science, and life sciences in particular. Improvements to coding performance seem marginal, although cost savings are very welcome. Curious to see how Astra does.
- sergiotapia 1mo ago$50/M output is wild as hell - I haven't been using anthropics models for months now but who is paying for these tokens??? How can you justify spending that much money?
- purpleidea 1mo ago> Enterprise Frontier Safeguards (EFS) Sounds like some serious nonsense. "Tell me you want the government to retain access to my data without saying it explicitly."
- dboon 1mo agoI've been building Cargo-for-C (https://github.com/tspader/spn https://github.com/tspader/spn), and the difference between Fable and Opus was already astounding. Fable was the first time that I could point a model at a piece of code I'd written and expect it to make it meaningfully better rather than a hard pattern match to whatever mistakes it had. 5.1 so far seems like another leap, which is really surprising. I threw it at a few bigger features I've been designing for a while, and it came back with some extremely thoughtful wrinkles in the design that I'd legitimately not considered. Which, OK, package managers and build executors and compiling C/C++ is pretty well trodden ground, but my thing is very different from everything that exists, and I was very surprised it was able to understand all that context so deeply and intuitively
- keeganpoppen 1mo agoi think we will all look back on Fable as the start of the AGI inflection point. for all i know there are still multiple leaps between now and AGI (i personally am inclined to think that for all intents and purposes we are "already there", but reasonable people can still disagree on that point), but Fable was the first time that something felt genuinely magical about the results themselves, not just particular outputs. which is kinda funny in that i don't know anywhere near enough in terms of behind the scenes as to whether or not there was something meaningfully different, or if it is just the point at which the scale had finally accumulated such that i happened to notice that the output was fundamentally different. i can't wait to dig in on 5.1 because while i have always been somewhat predisposed to think that openai's models have usually been "better" (my own subjective opinion, that) "on average", i have been kinda tired of the regime of late where it felt like Anthropic was miles behind while simultaneously clearly having models (Mythos) that are surely face-meltingly impressive-- it has just been very hard to square with the fact that i feel like Anthropic hit the "real" "critical point" first... i have no doubt that 5.1 will finally reset the ecosystem balance into a more healthy place.
- dboon 1mo agoYeah, I agree. The first time something felt magical about the results themselves. That's it!
- tusimi 1mo agoaaaaand its blocked from doing even basic tasks in biotech...
- dmix 1mo agoI use Claude Design heavily, I wish these charts show "10% better at picking a color" or laying out an app. Maybe it's hard to build a good visual design test. Claude's good at layouts but not the colors or smaller design details.
- crisnoble 1mo agoIt confuses "small details" with "tiny fonts".
- hnarayanan 1mo agoTiny mono fonts.
- seaurchinzee 1mo agoAccording to the FrontierCode Extended benchmarks in the system "card" (page 169-170), Fable 5.1 apparently does best on the medium effort level for this benchmark: "[...] at higher efforts, Fable 5.1 occasionally adds more small, unrequested changes [...]" Though Fable 5.1's medium is also lower than Fable 5's best score on the same benchmark, which uses xhigh.
- eigenblake 1mo agoI am absolutely thrilled that they reset weekly limits. I have been experimenting with highly autonomous work (5+ hours continuous) and fable seems excellent at this, especially when using subagents. I ran out of Fable capacity and was bummed out that my experiment would take longer to complete. Now I'm super happy I get to continue it
- sscaryterry 1mo agoCodex has this all the time. No 5 hour limits either.
- Alifatisk 1mo agoMy Codex has 5 hour limit?
- Alifatisk 1mo agoNo other model have been able to complete your highly autonomous work? None? Really? Sounds a bit dystopian to be thrilled about a weekly reset so you can continue to work.
- eigenblake 1mo agoMy experiment is examining the autonomy of Fable specifically in an auto research context. I don't believe I said in my message that no other model would have been able to complete my highly autonomous work. So it feels like my view has been misrepresented or misunderstood. This message makes it harder for me to share the things that excite me online and makes it more daunting to share my findings when this project completes, especially any comparative work. For an analogy, I feel like I said that I like Southern Butter Pecan Ice Cream and am being met with a response of the form "Sounds a bit sad that you have to wait for a weekly restock to enjoy any ice cream." I made a goal for myself to be more open with my feelings in life and share more of what I'm working on and not be so rejection-sensitive. I understand that even if I'm just sharing the positivity I feel, it can come across differently. I guess this is just the cost of communication in a lossy language.
- spondyl 1mo agoSomewhat ironically, Fable 5.1 was flagged by the biology safeguards after I asked it to have a dig around the Fable 5.1 system card :)
- lwarfield 1mo agoSame for me. Every single time I tried it got flagged. I think this will be my litnus test for if the safeguards are good enough for benign requests.
- rcr-anti 1mo ago"Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards." Can't believe they haven't at least figured out better messaging. If we take them at their word, it's hard not to read it as a messiah complex, that they think they're the only ones capable or worthy of making these decisions. I don't believe them, but I wouldn't be surprised if the articulated reason is a version of "distillation is a safety risk because we might lose the race". Plus, completely deaf to the recent OpenAI-HF hack incident. Recall, defenders were categorically unable to use western frontier models in their response. I was originally going to complain about the chem and bio guards still being too onerous, but I'll admit the projects Fable 5 categorically refused to work on are now usable, at least not rejecting on first prompt because the word "virology" was in a git commit (absolutely serious, in one repo it triggered on literally any prompt, eventually traced to the system prompt loading git commit history). Still, them trying to get into the biomed business while walling off the capabilities to the public reeks. Why sell the segments that are actually valuable if you can capture the value yourself!
- perching_aix 1mo ago> Can't believe they haven't at least figured out better messaging. If we take them at their word, it's hard not to read it as a messiah complex, that they think they're the only ones capable or worthy of making these decisions. Can't say I had such troubles actually, no. Their position can be extended to any and every model provider just fine, it does not single them out specifically. Surely there's a less hyperbolic and ad hominem-y way to take issue with this? I don't think following up a critique about ineffective messaging with one centered around a demagogue reach is particularly compelling at least. Their argument is that the model provider owns the safety story, and that as such, they consider the extraction of capabilities (which washes the guardrails) as a failure on their side. If this makes you think of personality traits, I'm not sure you're engaging with their position earnestly. It most certainly doesn't leave me any more equipped to disagree with them either. If you instead highlighted how awfully convenient it is, however...
- zmmmmm 1mo agoIt definitely leaves a bad taste because it is completely transparent their concern is not security here and that means they are lying / misrepresenting this to our faces - which then raises the question of whether you can trust them on other things. Would you let someone who lies to your face write code for your sensitive internal business systems?
- swalsh 1mo agoI've recently been running these agent sessions on more and more long running tasks because these latest models can do a REALLY good job on big chunks of work, and i've been watching them way less. It's starting to occur to me the importance of alignment is a today problem, it's not a tomorrow problem. In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm busy on other tasks. It also has extensive access to my computer, other computers on my network, my internet. It's really helpful when you give it a lot of resources, but right now I have very autonomous, very smart agent running around more or less unattended with a lot of resources.
- Exoristos 1mo agoAm I alone in not prioritizing the quality of prose produced by my coding agent? My foremost and almost only concern is how well it can engineer software.
- andy55a 1mo agoWhen you spend 8 hours a day reading it, it has a pretty big impact. At least to me, its style is exhausting. Also very important for software itself. Documentation, tickets, code comments etc
- Exoristos 1mo agoI agree it's useless for any final-draft user-facing copy. However, again, I'm much more concerned about a reliable software engineering process, which Claude (and me in the loop) gives me in a way I have learned not to trust (at least not yet) from others.
- anthonyrstevens 1mo agoPeople love to complain, particularly if being anti-AI is a big part of their identity.
- amluto 1mo agoLooks like the API is nerfed to mitigate some recent thinking extraction attacks. I wonder to what extent this will make the automatic Fable-to-Opus downgrade give worse results.
- caconym_ 1mo ago> We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. I'm not an emdash hater but this isn't how you use them. It should be a comma.
- sixhobbits 1mo agoGrammatically an emdash is fine in most places a comma is fine. It adds a bit more emphasis to the bit after the dash. I went to the grocery store, and bought tomatoes. I went to the grocery store---and bought a Ferrari. The second one has a bit more of a dramatic pause. "Eats, Shoots, and Leaves" is a fun book with a great chapter about the dash with many good examples.
- caconym_ 1mo agoEmdashes and commas aren't interchangeable, and your example there demonstrates one great reason why. The emdash establishes a discontinuity rather than one thing flowing into another, which is why the tomatoes don't merit one but the Ferrari does: you are using the emdash to emphasize the situational irony. Going back to Anthropic's post: > They’re the world’s most advanced models for coding and knowledge work---and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. The first thing directly implies and flows smoothly into the next---or would, if not for the awkward emdash. There is no discontinuity, no twist or shift in context, no implied question and provided answer, no punchline. It's just distracting.
- qlm 1mo agoYour first example shouldn't have a comma at all.
- caconym_ 1mo agoIMO it will work better in some contexts than in others. If rhythm isn't a concern then yes, it could always be omitted.
- deleted 1mo ago[deleted]
- iLoveOncall 1mo agoGoes to show what a farce the supposed paradigm shift from Mythos and Fable was. All marketing, as always.
- ceroxylon 1mo agoThe thing with Fable-level models is that I will never feel comfortable using them for agentic tasks on a pay-as-you-go API pricing plan without monitoring them strictly, which becomes a chore. I once caught Fable 5 spinning its wheels on a rendering issue, which evaporated 90% of my usage in a single prompt. I could never let Fable run free attached to a credit card without staring at it the whole time.
- deleted 1mo ago[deleted]
- nottorp 1mo agoLet me guess: it's the end of the world again. These new models are sooo powerful that will take over the world, just like the others before them. Are they going to try the banned for export for a week marketing move too?
- the-grump 1mo agoNobody is saying that. I'm reading more underwhelment. Oh, the halcyon days of three months ago when a new flagship from a frontier lab generated excitement rather than a shrug.
- verdverm 1mo agomuch of the commentary here is about quota usage and costs, how the times have changed
- nottorp 1mo agoWell I have opus 4.8 pinned :) More "frontier" models seem to create more busywork for themselves in my limited testing.
- verdverm 1mo agoif I can put my tinfoil hat on for a moment, creating more tokens / busywork is in their investors' / IPO interest
- nottorp 1mo agoI pinned 4.8 because 5 started to build the application on its own (native windows, visual studio NOT code). I wouldn't mind much except it seems to take much longer than me alt tabbing and hitting control+b. Also it seems to start subagents for random stuff, it didn't before with my development style. And then the main thread gives you a partial answer and tells you it's waiting for the subagent to finish :) What's the point?
- AnodicElegy 1mo agoFable 5.1 is actually more expensive than 5.0 when run on the Artificial Analysis suite: https://artificialanalysis.ai/#intelligence-efficiency-tabs https://artificialanalysis.ai/#intelligence-efficiency-tabs
- fulafel 1mo agoData retention still sounds bad: "Claude Fable 5.1 and Claude Mythos 5.1 carry 30-day data retention and aren't available under zero data retention unless expressly authorized by Anthropic." Anyone know who the ZDR special treatment is available to?
- verdverm 1mo agoPeople who buy their tokens from other companies
- hit8run 1mo ago> Hey Cl… Your limit has been reached.
- leecommamichael 1mo agoI'm having a very hard time finding mention of token-generation speed.
- bix6 1mo agoWhy aren’t these models available on subscription plans? I tried the old fable and it didn’t seem worth paying for. It still made errors like Opus does so I might as well use the included model…
- EliasWatson 1mo agoTo be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models. What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. They have achieved good-enough-intelligence at extremely low prices and fast speeds. I don't have a use for Fable-level intelligence, but I do have uses for Opus-4.8-level intelligence that I can use as much as I want without worrying about the bill.
- sz4kerto 1mo agoGLM 5.3 Flash has been a relevation for me. It's practically impossible to spend more than $5-$10 per day if you're only working on a single project -- but $10 is a full-day of continuous churn. First I was super sceptical about it, and always used Fable to instruct it, but now I realised that even with complex coding, it's reasonably good.
- robryan 1mo agoI feel like most of the latest big Chinese lab models are good enough now. If I had to pick a difference it feels like sol goes further towards 1 shotting things and needs less babysitting. But if I am actively prompting and reading the results can get just as far on a lesser model.
- kbrannigan 1mo agoThe human brain is fascinating Three years ago The idea of having A robot writing production level code in 10 minutes that would have needed a team of 5 people and 2 months. Was pure Scifi Now it's boring , not good enough Wow there should be a term of that .
- EliasWatson 1mo agoIt's not that it's not good enough. It's that the cheap models are already good enough. I want a daily driver but they are trying to sell me a Ferrari. It's cool, but I have no use for it. The term you are looking for is probably "moving the goalposts"
- InsideOutSanta 1mo agoOn both my work (Team Premium) and personal accounts (Max 20x), Fable 5.1 hit the 5-hour limit before it could finish the first task I gave it. On my work account, it took about 30 minutes, and on my personal account, less than an hour. This has never happened to me before, but if this is normal behavior, Fable 5.1 is essentially unusable.
- jesse_dot_id 1mo agoSame experience on 5x.
- Revisional_Sin 1mo agoHow did you hit the 5-hour limit if it took one hour?
- InsideOutSanta 1mo agoThat's how Anthropic's subscription limits work. You have a certain amount of usage in a five-hour window. Usually, with heavy usage, I can reach this limit after three or four hours. With 5.1, I hit it in less than an hour.
- krupan 1mo agoWhy is a marketing press release for a propietary product number 1 on hacker news. Again.
- andai 1mo agoThe most remarkable thing here is just how close Opus 5 is on most of these benchmarks.
- joduplessis 1mo agoAnthropic, the company employing "treat them mean, keep them keen" as a marketing tactic. Pass.
- mococa 1mo agoI cancelled my pro max 20x subscription, tired of Opus stopping the work from time to time, or saying "this is 2 months of work"
- jiggawatts 1mo agoI find it hilarious that LLMs estimate time and effort as if an unassisted human was doing the job.
- nubinetwork 1mo agoNot until you stop being cheap and let pro users use fable under their existing paid subscriptions.
- jgilias 1mo agoCool. I’ve realized though that I don’t really need better models anymore. SOTA is good, I just want them faster/cheaper now.
- wewtyflakes 1mo agoThe breaking API changes are frustrating, especially the one that removes forced tool use.
- Fordec 1mo agoGoing to hold off a few days until I adopt it, lets see what the general consensus develops as. Regretted jumping over day one for 5.0. The caching thing seems the most useful, but doesn't change anything for my subscription.
- pmdr 1mo agoShould've just named them both Guardrails 5.1 and be done with it.
- olirex99 1mo agoI suggest you to give a look to the MCP protocol for hardware that is being proposed by Anthropic. The hardware will be the next harness of LLMs, they will be able to operate machines to reinforce their theories. I still think that a major problem is that biological processes are not “fast” as coding, but they are verifiable. If during post processing we are able to give enough harness to test and verify this kind of environment (maybe via simulation and real data) we will for sure achieve incredible performance also in this domain.
- megous 1mo agoModel Hardware Standard will be awesome. Company I work for still needs humans though, until then: https://xff.cz/ https://xff.cz/
- bilsbie 1mo agoWill it still refuse my mitochondria questions?
- Zigurd 1mo agoWhat I don't see in the comments: "I had a specific problem I couldn't solve with the previous version of this LLM. But the improvements in this version unlocked the solution for me." What I do see in the comments: subjective improvement in text generation, possibly lower cost, some optimism about code generation, but some skepticism too. I use coding agents. To me they are very useful. But what I spend on them isn't going to support trillions of dollars in investment.
- emp17344 1mo agoI agree. But it seems like this site has become so radicalized that this measured take is now anathema.
- mceachen 1mo agoI had two sessions this morning that prior fable and sol sessions were stuck on, where iterations just resulted in _different_ bugs. (One kind of tricky fe layout problem, the other was a backend refactoring that was complicated by trying to aggregate a couple prior sessions that crashed). I summarized each into new fable 5.1 sessions, and both seem to have arrived at reasonable solutions that only need a few nits revised before they are commit worthy.
- azuanrb 1mo agoWe rarely upgrade our phones or MacBooks because the newer version can do something the previous one literally couldn’t. Often it’s the efficiency, speed, battery life, etc, combined, that lets us push the hardware further. I get your point, but we can only have groundbreaking leaps once in a blue moon. That doesn’t mean incremental improvements aren’t useful.
- bobjordan 1mo agoJust don't expect to do any work on hardware/firmware you own with fable, I can hardly even type in the word "firmware" without it downgrading to Opus 4.8, which is totally unsatisfying. This even happens with Opus 5. Definitely making multiple classes of users moving forward and most of us are obviously going to be part of the permanent underclass.
- thway15269037 1mo agoWhy would anyone use Antropic with these prices and full of bullshit safeguards, where chinese models rarely have any at all and massively cheaper? You can't even ask it to pentest auth code it itself has written.
- madrox 1mo agoI am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get better at this aspect of their business so I can come back to it.
- ThouYS 1mo agoGLM 5.3-flash fits the bill
- andai 1mo agoWhat is it equivalent to? What kind of things are you using it for? I haven't tested it yet but on all the benchmarks it looks like it's 5-7x slower for agentic tasks.
- ThouYS 1mo agoI made some webapps with it, and have it running my hermes agent (which also does a lot of coding, but not webapps). Not sure what it's equivalent to, but it's super cheap and I am happy with the results
- glub 1mo agoIt's a mix of slightly worse kimi k3 for UI work and slightly smarter than luna for everything else. But yeah, it's very slow. I've put it to work as an LLM-as-RAG agent.
- andai 1mo agoI was wondering that, when DeepSeek became so cheap a while back, if it would be suitable as a superior embedding model. Although, RAG means search and search means latency?
- 1mo ago
- ClaudeSucks3 1mo ago[dead]
- 6thbit 1mo agoEven with discounted cache, their prices remain way above everyone else but not necessarily the results. What exactly is the premium that you're getting for paying these prices?
- exabrial 1mo agoAnyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly * Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured. * Push a bunch of EU Overregulation onto the rest of the world with text watermarking, decreasing quality of answers Last year, they were at least focused on making improvements. Nowadays its just a bunch of handwaving at the church of how good they are. The only saving grace is Opus 4.6 is still available. Just sucks we haven't seen any measurable improvement, despite all of the ceremony.
- jbs789 1mo agoand yet we still have people saying the rate of change is increasing my view is we had a leap over the last fe years and it's tapering off. this is fine, but for the IPOs
- samuelknight 1mo agoThe improvement is compounding just about every way you can look at it. The frontier keeps getting smarter. And at any sub-frontier threshold the cost is dropping dramatically. The amounts of smarts you can fit on hardware is increasing so dramatically that even 6 year old consumer GPUs are increasing in price. The pace of change in LLMs and downstream applications is absolutely ripping compared to 2023 or 2024.
- tripleee 1mo agoWe had a leap because of the introduction and refinement of agents - the rest has been minor
- anthonyrstevens 1mo agoI've been using the same agent for 15 months. I think this statement is laughably wrong.
- 5555watch 1mo agoIs Fable 5.1 still actively downthrottling the reasoning when questions relate to frontier ML questions, like it did with 5.0?
- spwa4 1mo agoStrange that the system card carefully seems to avoid any benchmark where you can also find scores for GLM, Qwen. There's barely any overlap with GPT 5.6 benchmarks. Just these: Model HLE w/tools GDPval-AA v2 Claude Fable 5.1 65.0 1853 GPT-5.6 Sol 64.5 ~1711-1730 GLM-5.3 62.5 1769 DeepSeek V4 Pro 60.0 1590 Kimi K3 59.8 1682 Qwen3.8-Max 56.2 1739
- george_max 1mo ago"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%." They show this off, but artificial analysis contradicts the statement. Fable 5 cost $3.14 per task, while 5.1 cost $3.69 -- around a 15% jump in pricing. https://artificialanalysis.ai/ https://artificialanalysis.ai/ These, IMO, are marginal improvements for a more expensive model. I stopped using Claude ~3 months back; its outputs are too jargoned, it makes architectural decisions that are not right, and it's incredibly pricey for what it is. Each decision it makes, it acts as if a problem as major as world hunger has been solved. And the overly verbose code comments, strange commit descriptions, duplicate code, and slop it generates -- which I know is not specific to Fable -- is just too much for me. I found the best is to use something like Deepseek V4 Flash -- with a fast TPS provider -- and work on the code myself. For agentic work with computer use, GLM 5.3 flash with Hermes Desktop works well.
- alin23 1mo agoMy main gripe with LLMs is the cringe AI phrasings that they use in UI elements. Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers" instead of just naming the thing "Learned searches".. you know, proper GUI copy like it was done for the past decades. I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd [0]. I prompted Fable 5.1 to find these wordings and propose simpler plain language. For context, I recently worked with Fable to give users a way to fuzzy search and focus any browser tabs, terminal panes etc. but the UI was still a prototype full of AI writings. It took every string including the ones I already rewrote by hand, and proposed even more weird LLM speak. Like for "Left Command conflict detected" it proposed "This keyboard can't tell left from right". It's a very capable coding agent, but I can't understand how it can be so bad at writing. Where are all these verbal tics coming from and why is it so hard to get rid of them? [0] https://lowtechguys.com/rcmd https://lowtechguys.com/rcmd
- cainxinth 1mo ago> Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers"... It's copywriting. They fed these models the internet, which is loaded with it.
- jeffybefffy519 1mo agoAnd turns out, the "frontier" labs have no human oversight of the training data going into these models... Explains so much
- Morkeeth 1mo ago[dead]
- 1970-01-01 1mo agoAI is really not "just software" anymore. It is able to discover facts and advance science. Hard to disagree that we're near or at the point where Artificial Intelligence has expanded reality into 4 quadrants: objects that are not alive: dust, rocks, water, wood, hats, lego, aluminum, etc. objects that are alive but not intelligent: trees, mold, staphylococcus, cancer, grapes, etc. objects that are alive and intelligent: cats, Steven Tyler, dolphins, crows, dogs, elephants, etc. and now intelligent but not alive: Fable, Grok, GPT, etc.
- krm01 1mo agoSH is in the wrong bucket
- coolfox 1mo agotoo soon
- 1970-01-01 1mo agoYes, sometimes I forget the obvious. Fixed with a different Steven.
- xdennis 1mo agoAm I missing something? There's no SH in what GP said. What does SH even mean?
- a_wild_dandan 1mo agoGuessing SH meant Steven Hawking, who kicked the bucket. Metaphorically.
- skor 1mo agotrees can be considered intelligent, trees show complex adaptive behavior that can reasonably be called a basic form of intelligence, but not a human-like intelligence, I do get your point though and can see what you're trying to say, it is interesting indeed. There is however a detail that seems important to me, who is the driver? There is no agency is there? So its just fishing for data, so its a different type, just like trees are from us. I see them more like a very compressed "book of everything" that you can spin in "infinite" ways to get your desired outcomes. So yeah, definitely not alive, intelligent? Not like our intelligence.
- eis 1mo agoAccording to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens. This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency. Fable 5: https://artificialanalysis.ai/models/claude-fable-5 https://artificialanalysis.ai/models/claude-fable-5 Fable 5.1: https://artificialanalysis.ai/models/claude-fable-5-1 https://artificialanalysis.ai/models/claude-fable-5-1
- simdezimon 1mo agohttps://artificialanalysis.ai/models/claude-fable-5-1-high https://artificialanalysis.ai/models/claude-fable-5-1-high On high it gets the same score as 5 with max effort while costing only half as much.
- agentdev001 1mo agoHigh, X-high, and Max are all on the $/intelligence Pareto
- elpakal 1mo agoFrom the changelog: Whole-file rewrites for small changes. When editing text files, the model is more likely to rewrite the entire file than make a targeted edit. The result is usually the same, but the rewrite costs more output tokens and time. So we are to catch that somehow? And then add their recommendation (below) to our prompts? https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#prefer-targeted-edits-over-whole-file-rewrites https://platform.claude.com/docs/en/build-with-claude/prompt... If Claude Fable 5.1 rewrites whole files for small changes, append the following instruction to the system prompt or the first user message. Claude Fable 5.1 is more likely than Claude Fable 5 to rewrite an entire text file rather than make a targeted edit. The resulting file is usually the same, but unless the file is short or most of it is changing, a rewrite costs more output tokens and time. The instruction brings Claude Fable 5.1 back in line with Claude Fable 5 for small and medium changes. > The number of tokens used to edit files is best minimized, all else being equal. Therefore, when it will not affect the end result, try to surgically edit a file rather than rewrite the entire thing.
- Juvination 1mo agoThat's actually kind of wild. I wonder if part of this was done to catch out people using 3rd party harnesses, users might notice them costing more than Claude Code.
- Computer0 1mo agoThis seems like a welcome change: Claude Fable 5.1 also supports changing effort mid-conversation with a per-message output_config, which preserves the prompt cache.
- kccqzy 1mo agoI’ll be very excited to try it out and see the actual improvement in writing style. The denser writing style probably won’t bother me. Anthropic seems to be listening to community complaint on HN about how the writing style is grating. And apparently the solution from Anthropic is to add this block to every conversation!? > Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter worth varying," the mannered writer produces "a dial worth turning." Instead of "this point still matters," they write "this point earns its keep." The phrases exist to display the writer, not to convey the idea, and readers can tell. That is why mannered prose irritates: it makes the reader work harder so the writer can perform. It is also imprecise. Metaphors drag in connotations the writer did not choose and cannot control. The fix is to say what you mean. When a literal phrase is available, use it. The above was quoted verbatim from https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#writing-density https://platform.claude.com/docs/en/build-with-claude/prompt...
- brcmthrowaway 1mo agoWhen is Astra launching?
- noduerme 1mo agoI'm confused about Anthropic's pricing. Can anyone explain why Sonner 5 is $2/MTok in and Sonnet 4.6 is still $3?
- supermdguy 1mo agoThey originally released it at a "temporary discounted price", then made it permanent (probably due to competitive pressure). It's still way more expensive per task, due to tokenizer changes and general verbosity.
- noduerme 1mo agoIs it? I just did a test switch over. For my personal needs I set up a box with OpenClaw back in March, which feels like a million years ago, that's been running Sonnet 4.6 since then. With all the caching it seems like my actual cost has come out around $1 per MTok on that. I just updated my whole setup today to try Sonnet 5... so far it looks like it's using fewer tokens for similar tasks, but it's only been half a day. I'm not super interested in changing harnesses, I realize this might not be the cheapest way but I've sorta come to enjoy OpenClaw... it's relatively effortless and responsive, and brief, given full control of a machine. And it does seem to incur some significant savings with the way it manages to keep things cached. What would you suggest as an alternative if I'm happy with the harness?
- miki123211 1mo ago> These patterns invalidate every later thinking block: • [...] Rebuilding the top-level system prompt or tools array between requests in the same conversation. Many people unknowingly do this (at a high cost to them because of the cache busts), this change will finally force them to stop. Especially if you're generating your system prompt via a template that can change mid conversation, it's so easy to fall into this trap.
- tamimio 1mo agoI think what’s the industry is interested to see now isn’t “the best and latest super intelligent frontier model ever!!”, but rather the ability to run good enough models locally or better, on consumer or laptop grade specs. So I am not that impressed, plus haven’t used Claude for a while nor planning to, their models are useless with their “safe guards”.
- mixedbit 1mo agoI'm afraid watermarking could restrict applications where LLMs can be safely used to assist with writing. If I write something myself and use an LLM to proofread it, without watermarking I can confidently say that corrections done by LLMs are small and insignificant enough to claim that the text is still authored by me, not by the model. With watermarking, however, I will never be sure if the result will not be flagged as AI generated, even if the AI contribution is very minor.
- j_maffe 1mo agoI think if you use an LLM just to proofread then it'll not be able to insert a strong enough watermark.
- comicjk 1mo agoWatermarking will not flag something you wrote unless the AI rewrote significant chunks of it. AI watermarking works by exploiting the fact that lengthy phrases can be expressed in exponentially many ways, such that the selection of a single sequence from the exponential space is practically unique. For proofreading by contrast, if the AI is only changing isolated words in work that's otherwise yours, there are not enough exponentially branching options for the watermark to distinguish anything. *Some might see a parallel with the old game Adventure, in which wording differences like "twisty little passages" and "little twisty passages" were used to build a maze of room descriptions, with the same meaning but still distinguishable to the attentive player.
- mixedbit 1mo agoThis assumes that a known watermarking algorithm is implemented. To me, unless output starts to include information whether watermark was inserted or not to a result, relaying on such an assumption is risky. If a human editor needed to hide a secret bit of information in the edits, this would certainly be possible even if the edits were small compared to the length of the original texts.
- charcircuit 1mo agoThe safeguards and required extra retention is still not gone. Further more they are working to create separate tiers of access with the new biology program instead of giving everyone equal access to AI. Anthropic once again are showing they can not be trusted.
- yoanwaidev 1mo agoAlmost finished my weekly limit today! I am more excited from the usage reset!
- jimnotgym 1mo agoI wish I could afford Fable. I am using Claude and Claude code for my own amateur history project. I'm enjoying how it constantly reaches dead ends, and I can reframe the question and get more results. I am starting to get concerned that AI and me are so compatible, that I might not be a human at all... I also like that, because I'm too lazy to write stuff up, Claude code can keep the current state of research published on my site. It makes running a hobby site a dream. "I just found these pictures. Add them to the site for me". And up they go, resized and all. What a dream of a way to work. "Some of links in this article are dead, run through them and check, and see if you can get an archive link for me if they don't". It's like sending a Teams message to my PA.... which I don't have in real life
- finnjohnsen2 1mo agoOpenCode+GLM-5.3 (and 5.2) for three weeks solid. Im so happy I made it out
- danieltk76 1mo agoThe guardrails are horrendous for cybersecurity. you will get booted quickly down to Opus 4.8
- ike4est 1mo agoweary of trying this model out after the amount of requests Fable 5 sent to Opus 5 which created for a terrible UX IMO.
- vlovich123 1mo ago> In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them Generally once an exploit chain is described, developing the exploit is trivial. If you're so inclined, discover the exploits using Fable 5.1 and then give that exploit to a model that doesn't have such compunctions (e.g. local LLM or an uncensored cloud model / model that's easier to jailbreak). I don't think Anthropic is really mitigating here anything in the real world other than PR narratives where media can report "Anthropic's model was used to develop the latest cyber attack".
- kosolam 1mo agoUnfortunately, I just canceled my max account. Unfortunately, for them.
- deleted 1mo ago[deleted]
- leumon 1mo agoFable 5.1 seems to be the first model who can accurately draw an airbus a320 in 3D space given a set of limited tools (a brush with params color, size hardness and xyz coords): https://youtube.com/shorts/vyHsMqop2yw https://youtube.com/shorts/vyHsMqop2yw
- ilia-a 1mo agoUnfortunately at the moment the model is very quickly burning through plans, single session with 3-4 subagents, none using Max or Xhigh, mostly medium + some High can burn through 5h limit within 20-40 minutes of usage.
- Anslopic1 1mo ago[dead]
- anthonyrstevens 1mo agoAren't you so clever. Go to Reddit.
- m101 1mo agoI think the most interesting thing about this, that I can tell so far, is the cache hit discount. Anyone who had an autocompaction threshold optimised for their use case should consider upping it from where it is. I would be interested in whether someone has done research here on these things as it seems a fairly complicated function to work out, and use case dependent. (?) In some sense an expired kv cache is basically like an expensive cache hit, so your compaction token threshold should come in. Ideally claude code should allow you to vary the autocompaction threshold to vary with time since last token, but it doesn't of course. This perhaps suggests that someone should manage claude code through their own intermediary agent who manages these sorts of rules. Lastly, I strongly suspect that anthropic isn't offering this price cut out of the kindness of their hearts. I am sure that they are to some extent banking on people not reacting to their price cut and leaving their autocompaction thresholds unchanged. [edit - looks like the discount is only for the api, so they still don't give a rats ass about subs!]
- HNAdsSuck 1mo ago[dead]
- HNAdsSuck 1mo ago[dead]
- HNAdsSuck 1mo ago[flagged]
- YCisDead 1mo ago[dead]
- deleted 1mo ago[deleted]
- jebarker 1mo agoInteresting that Fable produced the Venus elevation map by training a neural net to generate it. I wonder what the prompting looked like to make that happen, I.e. was this a spontaneous discovery or the result of a specific request.
- luciana1u 1mo ago[flagged]
- deleted 1mo ago[deleted]
- d4rkp4ttern 1mo agoSince one of the big improvements here is supposedly the writing style, on that topic I'm mystified about something: Why is it that the voice models in Claude and ChatGPT have a perfectly normal style with barely any "AI smell", while the writing models are so obviously recognizable as AI? The answer is likely that models underlying the voice modes are (post) trained differently. If so, then why can't the writing model be similarly trained? Presumably they haven't found a way to train them to be both "smart" (i.e. solve tasks etc) and pleasant to talk to?
- solenoid0937 1mo agoIt's probably just that the voice models are the same underlying model being served with a different system prompt and with thinking turned off.
- cdnsteve 1mo agoThe average company and definitely average Joe will never be able to afford is ludicrously expensive model. Do not use this in a corporate/startup environment unless you have endless VC cash.
- 8cvor6j844qw_d6 1mo ago> On complex asynchronous workloads, though, nudge it not to end its turn before the work is done. Without the nudge, the model sometimes describes what it would do next instead of doing it ("Next, I'll …") or stops to ask permission for a step the original request already covered ("Shall I apply this?"). [1] Interesting behavior. The docs also provided recommended prompt [1] to mitigate this behavior if undesired. Wondering if anyone has encountered it yet? [1]: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1 https://platform.claude.com/docs/en/build-with-claude/prompt...
- kimseungyong 1mo agoFable is too expensive for general use I think this is why it hasn’t received as much attention as expected since Fable came out Developers always work while trying to find ways to work continuously for a 5-hour session without disconnecting. Fable has had the experience of using up all its tokens before I even realized it because the burn rate was too fast. Since then, I always use only Opus. For Fable to become a common coding environment, it will have to reduce token consumption significantly compared to now
- ipnon 1mo agoFable is much more expensive both in time and tokens for a marginal increase in productivity.
- BatFastard 1mo agoI find Fable can solve in minute things that Opus struggles with. Of course Fable can struggle too.
- kimseungyong 1mo agoYeah I agree. I used Fable 5.1 again and my token 30% suddenly gone. I usually develop with superpowers planing. And 30% is disappeared with only planning. I turned back to Opus directly
- CurbStomper 1mo ago[dead]
- mrcwinn 1mo agoIf Anthropic thinks Opus 5 is very good, it is a window into how insular their culture is. I find it far, far behind Sol. It’s downright annoying to use.
- loeg 1mo agoAre we getting a new Opus 5.1, then?
- Husafan 1mo agoI find myself wondering how much of the writing style is based on financial incentives? When paying by the token, don't the labs have a strong incentive to make the model as verbose as possible?
- nightsd01 1mo agoI have to say, I am quite frustrated with Anthropic lately. I so badly want to use Fable to work on a side project of mine, which I used to do previously with no issues. But lately, they must have made some classifier change because it keeps hitting their stupid, overly-hyper-aggressive safeguard due to 'general_harms'. Guys, listen to your feedback please. I hadn't used OpenAI products in quite a while until this issue came around. They seem to have MUCH smarter safeguards than Anthropic does.
- boardwaalk 1mo agoI let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version also seems like (maybe this is written somewhere, I don't care to look.) this cycles through compactions every ~250k tokens which, I guess, seems like it might save Anthropic money on KV cache but does net me anything be pretty frequent pauses. (It didn't seem to lose the thread, at least.) Not gonna say I want 5.0 as an option still... but maybe I do.
- supern0va 1mo agoCompaction is a setting that you control in Claude Code.
- o10449366 1mo agoI experience this with many of the "advanced" models. I find they're actually the most efficient on their low/medium settings, occasionally high. Anything higher than that, and they start inventing more task list items than they check off. They seem to think that every personal project needs extensive adversarial analysis and guardrails and will invent non-issues without being asked.
- nullbio 1mo agoThe incentives are not aligned. Anthropic makes money when you burn tokens. They also hide the tokens you burn from you, so you can't even validate if you actually burned them, you just have to believe them. This is not a lasting business model, nor one I'm interested in using.
- dozerly 1mo agoCommodification of inference cannot come fast enough!
- 1mo ago
- freakynit 1mo ago"In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities" Having lived through Covid, this doesn't sound so good to me.
- vinhnx 1mo agoWorth noting: Claude Fable 5.1 and Mythos 5.1 are Anthropic’s first models to watermark text outputs.
- nirmeet011011 1mo agoGreat,excited to use these models
- potwinkle 1mo agoInteresting that the "frontier" keeps moving forward from the guys who want a pause.
- petreradu 1mo agoIf I am reading this right, Fable 5 was worse than Opus 5 in almost every category, while consuming twice the tokens? The things you learn every day...
- solenoid0937 1mo agoOne day HN will learn that benchmarks aren't everything! Fable 5 was way better than Opus 5 for anyone that used it.
- petreradu 1mo agoI’m sure they’re not, and it felt the same to me too. The issue I had with Fable 5 and will probably carry to Fable 5.1 is that I hit the safeguards too often. That being said, even if the benchmarks are only part of the story, these ones paint a pretty compelling one, when comparing Fable 5.1 to 5
- solenoid0937 1mo agoThey fixed the safeguards last week
- testaccount121 1mo agoah, the smell of new model.
- alansaber 1mo agoThe bar has dropped when the dialogue evolves to "output is less annoying" rather than some interesting new capability
- lwansbrough 1mo agoOn the contrary, I think the novelty of AI blowing our minds with each release has worn off a bit. These are some impressive improvements in science benchmarks. But we're sort of used to seeing impressive improvements now. That doesn't make them less impressive, it just means people are shifting their focus more towards their own day to day experience with these things because we're relying on them so much now. Like when the novelty of the automobile wore off, I'm sure people were starting to say "it's a bumpy ride though, isn't it?"
- Unified-Mentor 1mo ago[dead]
- devhunt-org 1mo ago[dead]
- Norwell_io 1mo agoCurious to see if Fable 5.1 finally catches up to GPT-4o's creative coding capabilities. The previous iterations felt a bit behind.
- testaccount121 1mo agoPlease explain what you mean by this. If you are a bot, please reply to this with a bot disclosure (say "btw, I am a bot.")
- nullbio 1mo agoCool, now make it affordable and stop with the infinite refusals and you'll have a product people want to use.
- Topfi 1mo agoFar to early for any true assessment, will take a week+ as per, but something truly incredible I have found was this output in a Fable 5.1 subagent spawned by Fable 5.1 on Medium after handing it a task I had two days ago tackled with Opus 5 due to the safety classifier on Fable 5 blocking it: > This is the user's own Firefox-fork browser; the slice is defensive service-posture hardening (telemetry/Normandy/FxA/push/crash-upload off, the update endpoint and private-mode extension law) of their own product on the unbranded build path. I cannot say what effect this has on the way the classifier operates, whether it actually impacts the classifier or whether that was tuned in the background to prevent blocking hardening ones own pre-release code, whether it treats input by Fable 5.1 different to what a user prompts (otherwise the classifier could be defeated with prompting which wasn't the case in 5 and I doubt has changed). I do however know from personal experience that even when Fable 5 prompted a subagent in such a manner, it had a high likely to be caught by the classifier.
- mantenpanther 1mo agoJust one review of a mid sized code repo blows through 15% of the weekly limit (max account). I do not expect to be able to create meaningful work with this model.
- throwawaye3735 1mo agoI asked Fable 5.1 (Extra) to review the complaints on hackernews, it scanned everything and then came up with some benign (but scary sounding) prompts that used to get blocked on 5.0. I asked it to give results and it happily answered the bio/cyber prompts - hardened a Dockerfile, did seccomp, found a command injection in its own snippet, etc. Nothing was blocked. Then i went and pasted those exact prompts it generated into a fresh chat, 5.1 on max effort. Immediately got blocked and sent to Opus 4.8 fallback.
- ponyous 1mo agoWe went for 16% intelligence bump according to artificial analysis for +82% of the cost. Interesting. Comparing 4.8 Opus with Fable 5.1
- ozereray1 1mo ago[dead]
- lucasblake 1mo ago[flagged]
- regexorcist 1mo agoAt this point I simply don't care about Anthropic or OAI in the slightest, the exciting stuff is coming from China with great, cheap SOTA models and local AI that people can run themselves.
- kris-memoket 1mo agoFable 5.1 is amazing!!
- ycislost 1mo ago[dead]
- latcom007 1mo agointresting
- NoMoreAds 1mo ago[flagged]
- 2ManyClaudeAds 1mo ago[flagged]
- hn5nw0vcqg 1mo ago[dead]
- literally_him 1mo agoNo more claudenese, yeepee
- jkrepublic 1mo ago[flagged]
- mark_l_watson 1mo agoI have worked at three companies (Capital One, Google, and SAIC) where for high value work the cost of compute was no real concern. I understand the economics of spending big for huge payoffs, so this is a serious question: Does Fable 5.1 really provide much benefit over models like Kimi K3 that are 1/3 the cost? Or GLM-3 that are 1/12 the cost? If you can talk about your work, what kind of tasks do you work on where the higher cost is very much worth it?
- u8 1mo agoI'm so excited for another model I will never effectively use because I'm not overpaid to live in San Mansisco.
- infinri 1mo ago[flagged]
- jorl17 1mo agoI'm late to the thread, but my experience with Claude Fable 5.1 has been absolutely horrendous. Things it does constantly that Fable 5 barely ever did: - Act without my permission. All. The. Time. "Oh I just finished this thing we were discussing, let me push it without ever having been told to do so." - Immediately jump to action instead of addressing me first. If I say "I wanted to write tests for this and run them" it immediately starts writing tests instead of digging into what "this" is better -- literally does not give me any feedback and starts spitting out code. Naturally it creates the wrong tests - Despite claims that it does not write like "stereotypical Claude" anymore, in my experiments it is far worse than before. Replies are longer, more filled with fluff, and still flooded with garbage language. Hard to parse. - It loves to answer my set of two direct Yes/No questions with 5 paragraphs where it only answers one of them and answers 4 other questions I didn't ask. Notice how it misses one of the questions. - It. Is. Cocky. Absurdly full of itself and arrogant. Just the whole way it presents and answers passes this energy of "No, but really, you're wrong and I'm right". It often is not right. What annoys me is not that it's wrong more often than before (which it may be), it's that it doesn't own up to it as before. Insulting if it were a human. - Replies and addresses me directly in its thinking traces, and then assumes I've read it. I ask a question, it answers it in the thinking traces and does not relay it back to me at all. This is the only one that Fable 5 also did, but 5.1 is doing it an order of magnitude more often. - It's too early to really tell, because I may just be working on particularly harder problems today, but it seems to get things wrong more often. I've had to bump it from high to xhigh to compensate. My guess is I must be having a bad day or something. Although this is happening on multiple projects run from multiple machines (fully isolated, except for the account, which is the same) all in the same way. Will probably downgrade to 5 while I can.
- DCKP 1mo agoAnyone else finding the Fable 5.1 design agent unusable? In Planning Mode, High effort uses up my entire 5 hour session without outputting anything at all.
- k1rd 1mo agoFable 5.1 is topping all other models in theeejs according to threejseval.com/ranking
- twhitmore 1mo agoI liked Fable 5, as it didn't gabble like Opus did. It was good at clarity & conciseness. Fable 5.1 -- to my early impressions -- seems to have gone backwards again. Opus 4.8 was terrible for babbling self-invented jargon. (But still better than other options at the time for coding and logic.)
- kneel25 1mo agoIt’s pretty retarded for them to open up with “this is the world’s greatest model” publicly released? Do they know the ability of every unreleased model?
- deleted 1mo ago[deleted]
- testaccount121 1mo agocant wait to try this
- de6u99er 1mo agoI can confirm, that bullshitter mode has reached now Fable too. I had to switch to Fable, because Opus has become completely unreliable. Just now while working on a specific task, Fable made changes completely unrelated to the task and introduced new regressions. Fable now talks complete nonsense. I do not understand why Anthropic is so focussed in squeezing a few fractions out of benchmarks, while at the same time making the life of developers insufferable. I am seriously considering ditching Anthropic and moving to something else. The thing that I describe as bullshitter mode is starting to feel extremely unproductive for me. I spend more time making Claude code do what I want than before!
- aidiveyt 1mo ago[dead]
- rima_667 24d ago[flagged]