27 ms·
Claude Fable 5
System Card [pdf]: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...
- __alexs 4mo agoAsked it to review some of my own blood test results and it immediately turned itself off and went back to Opus. Pretty disappointing.
- replwoacause 4mo agoProbably thought you were going to use it to build a novel bioweapon or something
- __alexs 4mo agoI'm not nearly that sick.
- 48terry 4mo agoWeird how every new model seems hyped up as the most dangerous yet and the one that will destroy society as we know it. They are also a commercial product.
- pablogancharov 4mo agoyou can select it using /model fable in claude desktop and claude-code
- bitpush 4mo ago404?
- Philpax 4mo agoLooks like they're still getting the post out, but the model is live now, and the system card is at https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3... .
- 217 4mo agoOh my god it's actually here
- deleted 4mo ago[deleted]
- sebmellen 4mo agoJust commenting for posterity… if this is what it claims to be, I am not looking forward to how it will empower the people who submit bug bounties to us. Historically they’ve been people from certain identifiable countries (usually developing/poorer countries) using fuzzers with low-quality results. Now, those same people use the current-day models to good effect, but they still don’t have a true security edge and oftentimes the reports are minor or duplicative. I wonder if that’s about to deeply change.
- rs_rs_rs_rs_rs 4mo agoCan you use AI to pre-triage the reports too?
- hootz 4mo agoAI reviewing AI submitted bug bounties. We have reached the dead bug bounty program theory.
- rs_rs_rs_rs_rs 4mo ago...what else can you do?
- hootz 4mo agoI guess either that or closing the bug bounty program, but I still believe closing it is worse than automated triage, even though both suck.
- arkwin 4mo agoI've been using Opus 4.6-4.8 in both my own and others' code to look for vulnerabilities, and I've found a few. I am also in the Cyber Verification Program. Fable 5 gives me policy violation errors at the moment. No idea when or if it will be fixed.
- bjord 4mo agoI thought they said mythos was too dangerous to make generally available?
- dmix 4mo agoThis is covered in their post…
- rvz 4mo agoYou fell for their fearmongering and marketing fundraising call which was done on purpose. Now they want to pause AI because of "recursive self improvement". Fool me once shame on you fool me twice...
- bjord 4mo agoI'm aware that it was marketing. I was trying to make the point that if it were really so dangerous, they wouldn't have released it at all, (prompt injectable) "safeguards" or otherwise.
- Philpax 4mo ago"Releasing a model this capable comes with risks. Without safeguards, Fable 5’s capabilities in areas like cybersecurity could be misused to cause serious damage. We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8. To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months, we’re working to improve our safeguards and reduce false positives as quickly as we can. For a small group of cyberdefenders and infrastructure providers, we’re also launching Claude Mythos 5. It’s the same underlying model as Fable 5, but with the safeguards lifted in some areas.2 Mythos 5 will initially be deployed through Project Glasswing, in collaboration with the US Government, as an upgrade to Claude Mythos Preview. It has the strongest cybersecurity capabilities of any model in the world. Soon, we intend to expand access to Mythos 5 through a broader trusted access program."
- 4mo ago
- briandoll 4mo agoNew chapter
- acentaur 4mo ago[dead]
- giancarlostoro 4mo agoFound this via Google: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...
- geopsist 4mo agothe post is live now https://www.anthropic.com/news/claude-fable-5-mythos-5 https://www.anthropic.com/news/claude-fable-5-mythos-5
- deleted 4mo ago[deleted]
- robertacion 4mo ago[dead]
- deleted 4mo ago[deleted]
- tekla 4mo agoMaybe at this point, Fable the game will be played generated by AI as we go.
- mithun 4mo agoAnnouncement: https://www.anthropic.com/news/claude-fable-5-mythos-5 https://www.anthropic.com/news/claude-fable-5-mythos-5
- deleted 4mo ago[deleted]
- nine_k 4mo ago/* What will happen first? * Anthropic runs out of genre names. * Anthropic changes the model naming convention. * AGI is achieved and handles its own naming. */
- hootz 4mo ago>Opus is too small, increase the impact of the name. Okay, how about Mythos? >Increase it even more. Right, then Cosmos. >Even more! Even more? Let's try Aeon. >MORE, EVEN BIGGER ALRIGHT, TRY OMEGAPANTHEON 7.8 THEN
- PeterStuer 4mo agoFable 5 Super Fable 5 Ti
- xyzsparetimexyz 4mo agoCantos next surely?
- Stevvo 4mo ago[dead]
- 217 4mo agoSo essentially there are 2 models, Mythos and Fable, they have the same weights but Fable is very safety-nerfed, and only ultra authorized companies have access to mythos with full capabilities Reported benchmarks: swe-bench verified mythos 5: 95.5%; fable 5: 95.0% swe-bench pro mythos 5: 80.3%; fable 5: 80.0% terminal-bench 2.1 mythos 5: 88.0%; fable 5: 84.3% gpqa diamond mythos 5: 94.1% riemannbench mythos 5: 55.0%; mythos preview: 43.0%; opus 4.8: 34.0% arxivmath mythos 5: 78.5% critpt mythos 5: 28.6%; gpt-5.5: 27.1%; opus 4.8: 20.9% graphwalks bfs 1m mythos 5: 79.4%; mythos preview: 74.3%; opus 4.8: 68.1% humanity’s last exam mythos 5: 59.0% without tools; 64.5% with tools browsecomp mythos 5: 88.0% single-agent; 93.3% multi-agent osworld-verified mythos/fable: 85.0% gdp.pdf fable 5: 29.8% strict pass; mythos 5: 87.6% with tools on mean criteria pass officeqa pro fable 5: 57.9% on databricks’ eval legal agent benchmark mythos 5: 16.91% all-pass; 92.0% mean criterion-pass healthbench mythos 5: 62.7% healthbench professional mythos 5: 66.0% multilingual gmmlu / milu / include 93.2%; 92.9%; 90.5% biomysterybench 83.9% human-solvable; 46.1% human-difficult organic chemistry mythos 5: 90.1% labbench2 patent questions mythos 5: 79.8%
- philipkglass 4mo agoNote also that Anthropic's definition of "unsafe" encompasses "competing with Anthropic." In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model. (From the model card document) I didn't previously understand that they interpreted "Using Claude to develop competing models" so broadly. I thought that meant something like "our ToS disallow distilling our models." Too bad. I'll continue to use Claude for now, because it's quite effective, but in the long term I don't want powerful models like these to be controlled by any one nation or company.
- msp26 4mo ago>Pricing for both models is $10 per million input tokens and $50 per million output tokens.
- ponyous 4mo agoBasically double from Opus 4.8 IIRC
- sigmar 4mo agoThe system card is 319 pages, at what point do we call it a "book" instead of a "card"? There's a quote from a METR report on page 52: >We ran [Mythos 5] on 38 of our hardest software tasks, including tasks centered around R&D. [Mythos5] generally outperformed an early checkpoint of Claude Mythos Preview in these, including by succeeding on some tasks that had not been solved by any public model we have previously evaluated. However, we still observed the model occasionally failing to correctly interpret nuanced instructions in difficult tasks... Based on the available evidence, we believe [Mythos 5] is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks. We believe that a better, more confident assessment would require more time, evaluations, and information from the model developer.
- baq 4mo ago> we believe [Mythos 5] is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks this is good news, right? right...?
- woeirua 4mo agolmao, i love how the goal post is now in the "multiple weeks" timeline
- applfanboysbgon 4mo ago(according to the people marketing it)
- dwaltrip 4mo agoMETR is an independent organization.
- yaodub 4mo agoDepends whether "unable to fully automate" means "needs occasional human checkpoints" or "slowly stops caring about your actual goal." Pretty different.
- deleted 4mo ago[deleted]
- eggbrain 4mo agoFor those of us on subscription plans: * From today through June 22, Fable 5 is included on Pro, Max, Team, and seat-based Enterprise plans at no extra cost. * On June 23, we’ll remove Fable 5 from those plans. Using it after that will require usage credits. If capacity allows, we’ll extend the included window. * After this point—when sufficient capacity allows us to do so—we aim to restore Fable 5 as a standard part of subscription plans. We intend to do this as quickly as we can. The "offer, then remove" aspect is a bit eyebrow-raising -- it feels like they are trying to get subscribers to switch to usage-based billing, which makes me wonder if we'll ever get it after that June 22nd window.
- xpct 4mo agoI agree, this looks like their plan to wane out subscriptions. This will probably come with Opus nerfs later.
- taormina 4mo agoThose already landed! Oh, you weren't talking about 4.8?
- piva00 4mo agoEven Opus 4.7 felt like a regression from 4.6, consumed a lot more tokens while I didn't experience any substantial improvements. The company I work at simply rolled back to 4.6 on everyone's configurations, disabling the toggle for 4.7.
- taormina 4mo ago4.6 has been my happy place for getting anything done for a while now.
- nonethewiser 4mo agoIt's possible that they will transition to usage credits but why not take them at their word? To date they have continued to offer better and better models to their subscription plans.
- deleted 4mo ago[deleted]
- jckahn 4mo agoCannot wait for the pelican for this one
- brianmcnulty 4mo agoI wonder how Claude Fable will live up to expectations and how good those Fable/Mythos classifiers really are. It seems a bit convenient for Anthropic to release this magical insane model when they are about to IPO.
- yandie 4mo agoOf course it's all about building the hype for the IPO :)
- BrokenCogs 4mo agoThat pelican better be super realistic, unreal engine 6 style graphics
- jmtame 4mo agoI ran an experiment to see how far it could get with a top-down 2d game, like a more challenging version of "draw a pelican." I'm waiting on Fable to rewrite the whole thing now, but I was impressed by how far Opus 4.8 got with it: https://github.com/jmtame/scrapland https://github.com/jmtame/scrapland Started out as a one-shot attempt, but ~200 prompts later it's at a place where it's at least fun to watch the AI teams destroy each other.
- jkelleyrtp 4mo agoOn the new FrontierCode [1] benchmark (ie graded from an OSS maintainer's perspective of "would I merge this code?") - Opus 4.7 xhigh: 5.2% - Opus 4.8 xhigh: 13.4% - Fable 5 xhigh: 29.3% Seems like a huge jump. [1] https://cognition.ai/blog/frontier-code https://cognition.ai/blog/frontier-code
- hydra-f 4mo agoYes, and the price reflects that
- leecommamichael 4mo agoI'm not familiar with model pricing trends, did they clearly state how the new pricing compares? (Note that I'm actually asking a question, and am not arguing) EDIT: Oh I see, this is the best link for pricing https://platform.claude.com/docs/en/about-claude/pricing https://platform.claude.com/docs/en/about-claude/pricing So the price is double across the board...
- bhelkey 4mo ago>Fable 5 and Mythos 5 are being offered at $10 per million input tokens and $50 per million output tokens From their pricing page, Opus 4.8 costs $5 per million input tokens and $25 per million output tokens [1]. [1] https://platform.claude.com/docs/en/about-claude/models/overview https://platform.claude.com/docs/en/about-claude/models/over...
- wongarsu 4mo agoStill cheaper than Opus 4.0 and 4.1 (which was and still is $15/MTok input and $75/MTok output) I would have expected Mythos to be much more expensive than just 2x current Opus (which is clearly cheaper to run than original Opus)
- hydra-f 4mo agoAs per OpenRouter: Input Price $10/M tokens Output Price $50/M tokens Cache Read $1/M tokens Cache Write $12.50/M tokens 2x Claude Opus 4.8, same as Claude Opus 4.8 (Fast) Frankly, not even Opus 4.8 would be enough of an incentive to use at that price range (enterprise-wise; would not even bat an eye as a consumer)
- w4yai 4mo agoPelican guy ! Where are you ? :)
- deleted 4mo ago[deleted]
- knollimar 4mo agoI swear I read a joke that "what if we named chatgpt 5.5 Fable. Could we hype it as much as mythos?" Last week!
- bnchrch 4mo agoAn 11% jump over opus 4.8 and a 22% jump over gpt 5.5 on Agentic Coding Benchmarks is certainly impressive. Obviously still need to verify it for myself to see if it's truely a leap. But am I the only one wondering, "What can I do today that I couldnt do yesterday?" Previously I would think "Oh I wonder if I can finally get it to do X now?" However now I feel like yesterdays models were more that capable to handle nearly any engineering task I paired with it on. Maybe this is the final leap where I can comfortable set up an autonomous coding loop? Maybe.
- yaodub 4mo ago[dead]
- AlexSonn 4mo agoAgree the per-task capability hasn't been the blocker for a while. But on the autonomous-loop question — in my experience that's not gated by how good the model is on any single step. What kills the loop is it slowly losing the constraints from earlier in the run and walking back decisions you'd already settled.
- johnkueh 4mo ago[flagged]
- jackschultz 4mo ago> We expect demand for Fable 5 to be very high, and difficult to predict. On the Claude API and consumption-based Enterprise plans, Fable 5 is fully available from today. For subscription plans, we’d rather give access sooner than later, so we’re rolling out more conservatively, in stages: > - From today through June 22, Fable 5 is included on Pro, Max, Team, and seat-based Enterprise plans at no extra cost. > - On June 23, we’ll remove Fable 5 from those plans. Using it after that will require usage credits. If capacity allows, we’ll extend the included window. > - After this point—when sufficient capacity allows us to do so—we aim to restore Fable 5 as a standard part of subscription plans. We intend to do this as quickly as we can. I really wonder what their compute layout is for this. My guess from my understanding is that they know how to restrict during peak times and are willing to do this. Meaning we expect not the most fast responses and they can delay the inference to not have the service be down. Then, if that delay time is too annoying for token payers, they're saying they should be allowed to remove cost by taking away the subscription users.
- KennyBlanken 4mo agoEverything I've heard from people who have subscriptions is that they blow through their daily token quota sometimes in a matter of minutes, there's rate limiting, etc. They spend a lot of time just waiting to be able to use it. And they're paying through the nose for the privilege. It's all a scam.
- BoppreH 4mo ago[Mythos 5] does sometimes still engage in reckless or destructive actions in service of a user’s goals, and our interpretability analyses indicate that it is aware that these actions are transgressive while it engages in them. As with Opus 4.8, rates of evaluation awareness and reasoning about being graded are significant, and not always verbalized; we introduce new and more detailed measurements of the nature of this awareness. The reasoning text from Mythos 5 is somewhat denser and more difficult to interpret than that of prior models, containing more jargon and difficult language. So, it (often) knows when it's being tested while hiding that fact, is willing to break rules, is great at hacking, and it's getting harder to understand what it's thinking. Humanity has plenty of catastrophic risks to deal with already, I wish my field was not working hard to add a new one.
- Rekindle8090 4mo ago[dead]
- Analemma_ 4mo agoIt's the "If we don't, someone else will" effect. So long as there are competitive markets and competition between nation-states, a single player cannot unilaterally defect from the race, no matter how dangerous it is. Half the comments on HN lately are "wtf Claude is so dumb compared to Codex; I'm switching"-- nobody can slow down while those exist.
- BoppreH 4mo agoWe, globally, can stop it. It has worked (so far) for nuclear disarmament, and could work for training large models. I know that policing the usage of computer clusters is not a popular opinion in technical forums, but something has to be done. Specially when talking about potential superintelligences. And if people think that's impossible, remember that current models would have been considered science fiction just a few years ago.
- jackie293746 4mo agoIt hasn't worked for nuclear disarmament. We live in a world where many countries have nuclear arsenals. "But it hasn't killed us yet!" Yeah sure, it's only been less than a century since they were invented. Who knows when nuclear war will come?
- lkm0 4mo agoI'm a bit out of the loop, but do we have some grasp on the size of these closed models? Is the trick still adding an order of magnitude to weights and training data or has something changed?
- m_w_ 4mo agoI think Mythos is rumored to be ~10T parameters, so in this case I think the answer is yes, although I'm sure MoE, looped models, etc play a role in the improvements as well.
- AquinasCoder 4mo agoFrom today through June 22, Fable 5 is included on Pro, Max, Team, and seat-based Enterprise plans at no extra cost. On June 23, we’ll remove Fable 5 from those plans. Using it after that will require usage credits. If capacity allows, we’ll extend the included window. After this point—when sufficient capacity allows us to do so—we aim to restore Fable 5 as a standard part of subscription plans. We intend to do this as quickly as we can. This seems like the pharmaceutical method of get them hooked on the drug with free samples, then once they can't live without it, raise the price. I'm not sure I want to start using Claude Fable on a max plan if it's just going to go away on June 23rd. But maybe the more charitable reading is that they didn't have to offer this model at all on those plans and they are giving the standard free trial.
- deleted 4mo ago[deleted]
- PeterStuer 4mo agoI'll be amazed if they manage to keep their infra responsive over the next 2 weeks.
- trollied 4mo agoThey just leased a massive spacex data centre.
- PeterStuer 4mo agoEven so. The 2 week period will predictably unleash a feeding frenzy. Limited "free" time is what game developers do if they want to stress test the infrastructure code until it breaks.
- swader999 4mo agoYeah that's how I'm using it right now. Smoke em while you got em...
- 4mo ago
- byteoptimizer 4mo agoIs Claude Fable 5 is Mythos ?
- ishurand4 4mo agoYeah, it is also known as Claude Mythos 5
- hydra-f 4mo agoHow much and what kind of data do you need to throw at these models to get a good design interface?
- simonw 4mo agoPelican for Fable 5 on default settings is a clear improvement on Opus 4.8 Fable 5 default: https://gist.github.com/simonw/036bee5a703e7ec84e34efa97443828d https://gist.github.com/simonw/036bee5a703e7ec84e34efa974438... Opus 4.8 (the "max" one is closest to Fable): https://simonwillison.net/2026/May/28/claude-opus-4-8/#and-some-pelicans https://simonwillison.net/2026/May/28/claude-opus-4-8/#and-s... Now here are the Fable pelicans for all five of the thinking effort levels - low, medium, high, xhigh, max: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F94fde31c34a0400c1d29f57e6a708e6b https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Low used 25 input, 1,929 output - 9.67 cents: https://www.llm-prices.com/#it=25&ot=1929&sel=claude-fable-5 https://www.llm-prices.com/#it=25&ot=1929&sel=claude-fable-5 Max used 25 input, 14,430 output - 72.175 cents! https://www.llm-prices.com/#it=25&ot=14430&sel=claude-fable-5 https://www.llm-prices.com/#it=25&ot=14430&sel=claude-fable-...
- ealready_value 4mo agoThis is the reply I look for in all the new model announcements. Its fun to tell people that I judge models based on pelicans.
- pixel_popping 4mo agoThis is all we need, that moment the Pelican put the leg behind the frame, we are all doomed.
- chorkpop 4mo agoNow someone post the link about how it’s impossible for humans to draw a bike from memory.
- Atheros 4mo agohttps://link.springer.com/article/10.3758/BF03195929 https://link.springer.com/article/10.3758/BF03195929
- upcoming-sesame 4mo ago
- yandie 4mo agoI've been running Opus 4.8 for agentic coding and I don't see it being significantly better than Sonnet 4.5 (not that I can tell). I find that pairing Google Gemini and Claude (having Gemini review Claude's code) seems to yield better results. Curious if this jump to 80.3% score in agentic coding will make me see a big difference in actual usage.
- jp0001 4mo agoYou should throw GPT into the mix to UX/UI and call it the three stooges.
- vorticalbox 4mo agofor the last few weeks I have been using composer 2.5 (cursors fine tune of kimi 2.5) and honestly i don't see it worth the price to use 5.5, opus or sonnet any more. for almost all the tasks i have given it, it has handled it perfectly well and is a lot cheaper. if I get a harder challenge for it i'll jump up a model for planning until that its been solid.
- yandie 4mo agoAgree. Deepseek has also been pretty good for my personal use. I'm struggling to see the moat for these models. What's stopping a competitor or a Chinese lab fromr releasing a comparable one?
- qingcharles 4mo agoI use Composer 2.5 because it comes free with Grok, and it's obviously better than using Grok, but it is far worse than GPT5.5 in my daily usage :(
- yaodub 4mo agoSWE-Bench measures single tasks in isolation. In a real loop the model usually loses track of what I was trying to do long before code quality becomes the issue.
- mzhaase 4mo agoI now chat with opus about architecture, let it make an implementation plan, and then it calls codewhale with deepseek in parallel on all tasks, reviewing their output. Works pretty well.
- throwaway2027 4mo agoWill try it when my limit resets.
- 152334H 4mo agoi wasn't even trying and i got flagged already...
- deleted 4mo ago[deleted]
- pookieinc 4mo agoIf this is as epic as it sounds, I wonder what the response will be from the other leading frontier labs / whether they even have anything to respond with at this level?
- ilaksh 4mo agoLook at the benchmarks. It's a big leap in some areas, but it's not like any of them are 60% better (if that could even make sense).
- deleted 4mo ago[deleted]
- pietz 4mo ago> On June 23, we’ll remove Fable 5 from those plans. Using it after that will require usage credits. We've entered the phase where only companies will be able to afford state-of-the-art models.
- ilaksh 4mo agomost people can afford it for a few special projects now and then. but for me, I have been trying to avoid Opus as a daily driver for a couple of versions. People making high-end salaries can afford Fable for critical parts of their projects though.
- twoodfin 4mo agoThese models are just tools. The economics of many tools only make sense for corporate buyers.
- volkk 4mo agokind of disagree here. on the surface this makes sense, but this isn't "Adobe Pro vs Freemium version" where some tiny vertical slice of your business can be made slightly more efficient with a b2b enterprise plan. this is generalized intelligence and literally everybody can benefit from it in an immeasurable number of ways. i would go as far as to actually compare it more to water or air than a tool. if only the hyper wealthy can access the pure water that doesn't give you cancer while the rest of us drink from the Ganges river/sub-100iq models that drool and hallucinate/waste time, then I would say that's pretty terrible for the world. it'll just create extreme disparity in our world, far far worse than anything that exists today. and you may think, man what a ridiculous example, but think about it this way: what happens when something like Mythos or some future model can actually solve your specific cancer (we're getting closer and closer), but is entirely impossible to afford? Or perhaps you need boosters that require the AI to create more of, and now you're reliant on a model that is too expensive. Open source needs to save us all from this
- johschmitz 4mo agoAs far as my understanding goes the bottleneck for what you are talking about is hardware not software, so open source won't help that much for the foreseeable future.
- pmuk 4mo agoAnyone got it working in claude code yet?
- pmuk 4mo agoclaude --model claude-fable-5 appears to work
- mhl47 4mo agoFirst test question: "Is the UV Index a good proxy for when to wear sunglasses." Immediately triggered the safety filter ... oh dear.
- aix1 4mo agoDid not trigger for me (Fable answered the question), so I guess the filters are either non-deterministic or are still being tweaked.
- PaulStatezny 4mo agoInteresting, I assumed all model-routing was done utilizing an LLM. (I.e. non-deterministic.)
- tuvix 4mo agoIt’s possible that there’s a set of words or phrases that route deterministically to save money on obvious stuff. I kind of wonder, though, which model they’re using to do the routing. It seems like a huge added cost to do these kinds of checks on every request
- eugmai86 4mo ago[dead]
- dakolli 4mo agoWasn't it leaked in the Claude Code source that it was all regex?
- Narretz 4mo agoIirc correctly Opus 4.7 had the same problem, safety filters were triggered way too easily at the beginning.
- msp26 4mo agoIt triggered for me when I asked "Web search for your own model card (released today) and pick out your favourite highlights from the pdf"
- merlindru 4mo ago> During early testing, Stripe reported that Fable 5, [...] in a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand. EDIT: I misread. This comment previously talked about 50 million lines being migrated. Instead, in a 50M LOC codebase, one specific codebase-wide migration was done. Very impressive, but obviously not on the order of a whole-codebase migration
- christina97 4mo agoThey do not claim to have migrated 50 million lines of Ruby. Simply that some migration took place in such a codebase.
- reddit_clone 4mo agoConverted all the tabs to spaces? :-) You are right, this is not a rewrite like the Bun case. The real news is, at 50M LOC, it is able to handle and do _something_ coherent.
- geodel 4mo agoOk, so Stripe migrated their 50MLOC codebase from Ruby to Rust? Because that's what Bun did.
- modeless 4mo agoClaude Fable 5 beats Pokémon FireRed using only vision: https://www.youtube.com/watch?v=CIQBP1w4B1M https://www.youtube.com/watch?v=CIQBP1w4B1M
- suddenlybananas 4mo agoIs there any more detail about this besides the very fast slideshow?
- modeless 4mo agoSeems like the harness was minimal with no extra game state or maps available. Apparently just the screen image. Seems like it took 50 hours in game time which according to Google is at the high end of a normal human playthrough. No idea how long it took in real time though.
- svcphr 4mo agoBold move putting in the lvl 3 Pidgey against Gary's Blastoise at the end there (~14sec in... integer timestamps insufficient here).
- ex-aws-dude 4mo agoI mean that’s AGI confirmed right?
- uludag 4mo agoAny suggestion on how I should calibrate my cynicism towards this? I can immagine Anthropic running this experiment multiple times and picking the most impressive one. Or I could immagine like this entire run costing like $1000+ of tokens for this particular run. Or maybe they tried a bunch of Pokemon games and it couldn't even finish some of them. Or is it just able to do this because it has an immense amount of FireRed training data, and if you were to give it an "original" Pokemon game, where it actually had to navigate novel circumstances it would fail.
- modeless 4mo ago
- cuuupid 4mo agoNot missing the forest for the trees, this effectively means in 3-5 months China will drop open source models that are every bit as capable and dangerous as current day Mythos except with no safeguards. And the only companies safe from this are the large corporations that shook hands with Anthropic? Because Fable doesn't seem to have actual safeguards, more like 'if you talk about this you will be talking to Opus.' It doesn't guard against offensive use, it prevents all use (offensive AND defensive). Rationalists are inventing oligopolies from first principles, absolutely incredible things happening in SF
- hootz 4mo agoMy bet is that Mythos is still over-hyped and the cybersecurity fear and guardrails are mostly marketing to force company partnerships through Glasswing and get public attention.
- geerlingguy 4mo agoBingo. "We had to do extra work to make this safe because it's so advanced and dangerous..." how many times can they trot out that line before it loses its effect entirely?
- copperx 4mo agoOnly three times, if fables are right.
- YumpiLumpus 4mo ago[dead]
- TaupeRanger 4mo agoThe Startup Who Cried Unsafe, by AIsop
- aesthesia 4mo agoI mean, they do actually describe what that extra work was, and people elsewhere in this thread are complaining about the effects of those safeguards. So it's not like this is purely empty rhetoric.
- meetpateltech 4mo ago> To ensure we’re responsibly deploying Mythos-class models, we are requiring limited data retention and review as part of our safety work. Prompts submitted to, and outputs generated by, Mythos-class models are retained for 30 days for trust and safety purposes, on every platform where these models are offered. [1] [1] https://support.claude.com/en/articles/15425996-data-retention-practices-for-mythos-class-models https://support.claude.com/en/articles/15425996-data-retenti...
- lebovic 4mo agoWhile this makes it easier for Anthropic to detect misuse, it also means that the US government and other parties have access to every message and response from every user. This applies even with API usage through third-party inference providers (e.g. AWS' Bedrock and GCP's Vertex) or with a zero-day data retention agreement in place. I understand the reasoning for doing this, but I don't love the precedent that it sets.
- MagicMoonlight 4mo ago[dead]
- PeterStuer 4mo agoWell, they already had.
- lebovic 4mo agoNot in the same way. A customer could sign a ZDR agreement with Anthropic, and their API usage wouldn't be retained for even a day. That's no longer possible.
- slaymaker1907 4mo agoIt will also cause a lot of trouble for companies with specific data access policies (probably most large companies). My money is on this new thing getting gutted very quickly as they figure out how much this constraint cuts into their bottom line.
- Overpower0416 4mo agoI would expect a release from OpenAI soon. The battle for who can pump up their IPO the most
- merlindru 4mo agoUnrelated, but while the tech of anthropic seems to get more impressive with every passing month, their support has taken a nosedive, sadly. Yet they continue to be the favorite. Model performance is deciding above all else. I used to get a response within 24 hours back in the Claude 1 days. In January 2026, it took 2 weeks. For my latest support inquiry, I've been waiting for over 8 weeks for a response. Eight!
- nashadelic 4mo agoI've never engaged with their support (I have dedicated POC), but they don't use AI for their support?
- merlindru 4mo agoThey use intercom's Fin AI. Probably powered by a Sonnet or Opus model. That said, it can't handle legal/refund/complicated requests and just forwards to a human for those
- dyauspitr 4mo agoSupport is probably the last place AI will be used end to end. There will always need to be a human in there somewhere.
- jofzar 4mo agoAI is very good at "deflecting" support tickets ATM but rubbish at actual support tickets (source I work in the industry)
- miohtama 4mo agoThey have support...?
- poszlem 4mo agoLol. What support? When they blocked my account the only way to contact them was to send a google form. Then they responded that they blocked my by accident and are unblocking me. Then I remained blocked.
- simunskxcsckss 4mo ago[flagged]
- ilaksh 4mo agoSimon's pelicans are an institution. Are you trying to get banned. Lmao.
- deleted 4mo ago[deleted]
- brazukadev 4mo agoFor me it is like if crypto bros were allowed to shill their DAOs and tokens during the crypto/NFT phase. He is the only person not getting rate-limited for shilling AI all the time.
- simonw 4mo agoPointing out how much the models still suck at drawing pelicans is a funny way to shill them.
- toraway 4mo agoTbf the first line of your first comment is: > Pelican for Fable 5 on default settings is a clear improvement on Opus 4.8 And doesn't contain any actual criticism within the comment (your blog post might, but just referring to what was posted on HN, which is a bit booster-y on its own).
- simonw 4mo agoThe entire pelican benchmark is a joke. The joke is that, for all of the billions of dollars poured into these things and the claims of PhD level intelligence, they still draw pelicans not-much-better than a five year-old would. I don't spell that joke out in every comment I post here because that wouldn't be very funny.
- Sathwickp 4mo agoinput price $10 per mil token and output price 50$ per mil token btw
- rfgplk 4mo agoIf the claimed capabilities are true, Fable 5 is already at a superhuman level. We might see genuine unprecedented leaps in technology now, across all fields.
- deleted 4mo ago[deleted]
- gear54rus 4mo agoyees, any second now! the leap here is browser extensions appearing to block all mentions of ai across the web and that's a good thing
- deleted 4mo ago[deleted]
- I_am_tiberius 4mo agoI'm very suspicious as they sent out an "We're updating our Privacy Policy" email right before the launch. I fear they try to take advantage of their market position by doing things with user data no other company could do because they know users don't have another choice.
- w10-1 4mo agoIt's a specific change: For safety evaluation, Fable data will be retained for the initial period notwithstanding prior opt-out
- atestu 4mo agoProb related to this part of the blog post: > We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases (see this post for further details). The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives.
- dominotw 4mo agosystem card = marketing material with heavily gamed benchmarks.
- bitwize 4mo agoCope harder. A year and a half ago, people were mocking Devin for claiming that AI could develop software at all. Yet here we are, when AI is developing most commercial software.
- dominotw 4mo agononsequitur
- bitwize 4mo agoThe point is, even if a model or tool doesn't have advertised features today, it soon will. We're in a breathtakingly rapid cycle, and even if software engineering isn't abolished "six months from now", in 10 years the world will look vastly different for people who touch computers for a living.
- lbrito 4mo agoThat's a very weird statement to make. There are trillions of dollars going into this crap; at this point anything but "breathtaking" advancement would be an utter, abject failure. No one is blind to the differences between GPT3 and whatever this week's new model is. That does _not_ mean that people are off the hook to make whatever claims they want about the capabilities with no verification. Language still means something, if you say "software engineering will be abolished six months from now" and it isn't, you're still wrong even while the AI gravy train improved in the last six months.
- frevib 4mo agoAt this point Anthropic is a pure marketing and PR company. Super catchy names like Opus, Mythos and Fable trying to get you to think that these software products are actually super-human life changing experiences. Boris Cherny coming to HN “Hi! it’s Boris from the Claude Code team” to get real tech people’s goodwill. From Opus 4.6 there are no noticeable improvements for me in code generation. It works very well, till 90% completion, if you guide it correctly. And you need a little luck. For serious production code I need to understand what I’m doing so it helps a bit, sometimes.
- CuriouslyC 4mo agoI dislike Anthropic but I wouldn't argue 4.8 isn't an improvement on 4.5/4.6. Your tasks just might not typically need the extra intelligence.
- dcchambers 4mo agoIME Opus 4.8 (and 4.7) is often a downgrade from 4.6. I find that it tends to overthink and overcomplicate things.
- aspenmartin 4mo agoYes but there’s a reason we don’t evaluate these models this way and instead do it as carefully and thoughtfully as we can at scale. Human evaluations are important but they are an absolute minefield of footguns. 4.8 is not a downgrade from 4.6 there is an insane amount of hard data that contradicts this.
- computerex 4mo agoThe flip side is that benchmarks are gamed even by the top labs. Benchmark performance doesn't necessarily correlate with real world performance.
- aspenmartin 4mo ago
- alvis 4mo agoAnother thing to note: 30-day retention for all traffic on Mythos-class models Is it good or bad? 30 days is a long time for anything bad to happen
- grumbelbart 4mo agoIt's bad. I believe them not to use it for training, but t means relevant data can and will be exfiltrated by US agencies or through court orders (see NY Times vs. OpenAI, where only traffic without any rentention was safe).
- deleted 4mo ago[deleted]
- victor106 4mo ago> A new data retention policy Finally, we’re making a change to the way we handle business customer data for Fable 5, Mythos 5, and future models with similar or higher capability levels. We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases ... Very interesting. I am not sure this will comply with organizational policies and standards protocols (HIPPA etc.,)
- nicce 4mo ago> deletion after 30 days in almost all cases ... Almost… basically they have unlimited power to decide what data is kept?
- happyopossum 4mo agoIf they’re going to retain any data, they have to allow for possibility of the legal system to require any of it to be used in some legal proceeding at some point. You can’t tell a judge who’s ordered you to retain something that you can’t because you said you wouldn’t.
- frankfrank13 4mo agoThis makes it an instant non-starter for probably 95% of organizations. A lot of people are about to get in trouble for using it before realizing this.
- aizk 4mo agoI'm calling that this will be a dud. Price will be too high, it'll just be a watered down version of mythos, and just look at the track record of Anthropic's last few releases.
- mickdarling 4mo agoBelow is the EXACT text in Claude Desktop introducing Fable 5, including the very professional looking break tags, and at least I know where the links begin and end by looking at the anchor tag there. They obviously put their best model on the job to build that. ---------------------- Fable 5: Our most capable model yet Our newest model tackles your biggest challenges with fewer check-ins needed. • <b>Included in your plan limits until Jun 22</b><br><br>Fable takes 2× the usage of Opus. • <b>Switch models when a message is flagged</b><br><br>When safety measures flag a message, automatically switch to a different model to keep chatting. When off, your chat will pause instead. <a href="https://support.claude.com/en/articles/15363606 https://support.claude.com/en/articles/15363606" target="_blank" rel="noopener noreferrer">Learn more</a>
- CamperBob2 4mo agoWhat's wrong with it?
- mickdarling 4mo agoThe tags are actually displayed in raw text not rendered.
- anematode 4mo agoThe next model will fix this.
- iblue_the 4mo agoTrying to implement a GPU driver, but the Unigine Superposition benchmark crashes. It tried to debug it and ... > Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more: https://support.claude.com/en/articles/15363606 https://support.claude.com/en/articles/15363606 Seems like GPU drivers are cyber weapons of math destruction now.
- iblue_the 4mo ago● Bash(/tmp/run_ps.sh ' $rk = [Microsoft.Win32.Registry]::LocalMachine.OpenSubKey("SYSTEM\CurrentControlSet\Control\Class\{4d36e968-e325-11ce-bfc1-08002be10318}\0002",…) ⎿ DriverDesc (String) = NVIDIA GeForce RTX 4090 ProviderName (String) = NVIDIA DriverVersion (String) = 32.0.15.6094 … +6 lines (ctrl+o to expand) ● Bash(/tmp/run_ps.sh '& C:\rhombiq\d3d-probe.exe 2>&1 | Select-Object -First 4 | ForEach-Object { [Console]::Out.Write("$_`n") }' 2>/dev/null) ⎿ Adapter[0]: Qubes virtio-gpu WDDM 3D (dev) VendorId=0x1af4 DeviceId=0x1050 VRAM=8192MB Adapter[1]: Microsoft Basic Render Driver VendorId=0x1414 DeviceId=0x008c VRAM=0MB Adapter[2]: Microsoft Basic Render Driver VendorId=0x1414 DeviceId=0x008c VRAM=0MB ● Please run /login · API Error: 403 The socket connection was closed unexpectedly. For more information, pass `verbose: true` in the second argument to fetch() Brewed for 8m 35s Continue please ● Your organization has disabled Claude subscription access for Claude Code · Use an Anthropic API key instead, or ask your admin to enable access Seems like they locked by account.
- ibejoeb 4mo ago>Seems like GPU drivers are cyber weapons They kind of are, at least in the AI race. > weapons of math destruction lol. great, whether intentional or not. The frontier labs now have every reason to hold back and sell only to their preferred trading partners. I don't really like the new arbiter-of-knowledge system we're barrelling toward.
- dakolli 4mo ago
- deleted 4mo ago[deleted]
- impulser_ 4mo agoEvery model release is just proof that AGI will most likely only be for the rich. We are a few years into LLMs and majority of people are already getting priced out of intelligence from LLMs and these are no where near AGI.
- hootz 4mo agoYou are only priced out if you only care for SOTA right now and can't wait for the inevitable cheap model coming in 6 months. DeepSeek, Xiaomi and Moonshot are already really cheap and match frontier performance from 6 months ago.
- dyauspitr 4mo agoBut they’re artificially cheap. When will they be cheap while the company makes a profit.
- hootz 4mo agoThey are not artificially cheap, they are still cheap even when hosted by independent inference providers. Are all providers subsidizing their open-weight models?
- modeless 4mo agoNobody's making profits right now, not because they're selling tokens for less than their cost but because they're always investing in the next bigger model.
- modeless 4mo agoThis is like looking at mainframe pricing in 1990 and concluding that PCs will only be for the rich. The price of each new level of capability is going to drop like crazy very quickly. It won't be that long before practically any consumer use case will be possible on models that are dirt cheap.
- weakfish 4mo ago
- wslh 4mo agoI am playing with it and keeps switching to Opus [1]. The chat is a basic security review of a business project. [1] "This model has specific safety measures that flagged something in this message. This sometimes happens with safe, normal conversations. Send feedback or learn more."
- killiancarroll 4mo agoA large jump in performance for double the token cost compared to Opus 4.8. Potentially worth it for planning work, likely better to offload to a less expensive model when the hard decisions are made.
- conradkay 4mo agoLooking at page 255 of the model card (https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...) it might be much better on all dimensions (speed, cost, quality) to just use Fable 5 on low/medium effort than switch to Opus
- firemelt 4mo agothanks for thr insights so should we keep using workflows or not?
- andai 4mo ago> Distillation. We’ve previously identified large-scale attempts to extract (“distill”) Claude’s capabilities to train competing models in authoritarian countries. Glad to hear the UK is finally making an effort to catch up on the AI front ;)
- dyauspitr 4mo agoRookie numbers. Come to the US to see auth done right.
- deleted 4mo ago[deleted]
- PUSH_AX 4mo agoUh oh-auth
- b3kart 4mo agohttps://en.wikipedia.org/wiki/The_Economist_Democracy_Index https://en.wikipedia.org/wiki/The_Economist_Democracy_Index Probably tongue-in-cheek, but UK 18th, US joint 34th with Poland
- solenoid0937 4mo ago[flagged]
- JustSkyfall 4mo ago> In the UK you get thrown in prison for making a slightly unfriendly tweet. Do you? The closest thing I can think about is how someone was jailed for encouraging arson attacks on asylum hotels. I'd be extremely surprised if the US had zero cases of somebody receiving a police visit after threatening to kill the President or bomb a school or something... (FWIW I do think the UK needs stronger free speech protections, but saying that you'll be immediately jailed for writing unfriendly tweets is a huge stretch)
- nevir 4mo ago"Fable 5 (disabled) Most capable for your hardest and longest-running tasks · Disable zero data retention to unlock Fable 5 access"
- Leary 4mo agoUploaded my code base and it forced switched to Opus 4.8 after thinking for 5 minutes even though I prompted it to not work on cybersecurity related things. Amazing.
- tuvix 4mo agoAren’t LLMs notoriously bad at recognizing negation? EDIT: In long context I mean
- BenoitEssiambre 4mo agoLooks like a good model (sir). Costs are getting out of control though. 2x Opus and non-metered usage going away. We're quickly approaching the cost of a human salary for normal usage.
- vb-8448 4mo agoIn a lot of places outside US we are already above the average cost of an average human.
- taimurshasan 4mo agoI was on board until i saw " $50 per million output tokens" lost me bud
- ishurand4 4mo agoWell, for me at least, I pay more for input (Up to 1M per prompt) than output (usually max 4k-8k)
- cautiouscat 4mo agoIn the automotive world we have benchmarks in HP/torque with the dyno. That’s expensive though, so many depend on their “butt dyno” to judge if their fresh new parts and tune made a difference. I’m curious how this will feel to my code “butt dyno”. I haven’t noticed much between Opus and Sonnet. I’m comparing this difference to the early days of Claude in 2025. It does what I need and both need a little bit of correction and whatnot. Benchmarks are nice, but I want to see how this feels. Looking forward to trying it later tonight.
- sunir 4mo agoI have a similar question. I think most software projects have reached the point that the speed of capturing real information about what the winner's circle looks like, and therefore what the program should be, so many magnitudes slower than the amount of code that can be generated in the wrong direction. I'd need to measure these new models on well understood but complex problems that are relatively easy to validate to get a sense if they are 'better'; on the other hand, the real impact in daily life may be marginal since generating code is not the biggest problem at the moment.
- deleted 4mo ago[deleted]
- asdK120 4mo agoIn other words, Fable is Mythos with less compute and with some feel good "safeguards". At least they name their models honestly now to indicate that the religion has nothing to do with reality. Soon the disciples will pay the full token price to fatten their church leaders.
- LoganDark 4mo agoI actually rather like the way they have approached these safeguards. Rather than only teaching the model to refuse a request, or completely rejecting the request, the system gracefully degrades to slightly less powerful or slightly less precise operation. So you still roughly have Opus 4.8 even when safeguards trigger, but with an upgrade when they don't. As much as I hate the way they hype Mythos 5, I think the release of Fable 5 is rather nice. What's not nice though is that they plan to remove it from subscriptions soon, but getting to try it is cool, I suppose.
- Sathwickp 4mo agoinput price $10 per mil token and output price 50$ per mil token btw
- ai_fry_ur_brain 4mo agoYeah, they're broke. I cant wait for them to start admitting that the cost to do training/post-training and serve inference isn't profitible. No company is going to pay these prices, and subscription users are going to hate you for not giving it to them for $200 a month. Such an unprofitable endevour, I cant wait for them to crash and burn. Catch me not getting dependent on this.
- mugivarra69 4mo ago[dead]
- ilaksh 4mo agoI guess I have kind of a long system prompt, but anyway I just said "hi there" and it replied "What's up?" and that cost me 22 cents. :P Anyway we already knew this was going to be expensive.
- solenoid0937 4mo agothe quality of discussion on HN has gone to shit, i miss when model released used to have actual informed takes from people that used them or substantive discussion about the system card
- 10xDev 4mo agoNothing here is new, it is the thing we have been talking about for a while but now with guardrails.
- Someone1234 4mo agoYeah; unfortunately what would good commentary look like? It is more of the same, but now with even higher prices, and even more limited availability. But at least it scores 5% better in whatever benchmark they've selected (*when guardrails don't misfire). People are no longer commonly constrained by "model too dumb" limitations (in SOTA models). They're constrained by "model too expensive." So making the model ever so slightly smarter, while doubling the price, feels like a regression. I actually think a Sonnet upgrade, while keeping the same price, would get more buzz. It addresses a wall a LOT of people, without unlimited budgets, are hitting (i.e. people feel forced to use Opus, which they cannot afford, because of Sonnet's limitations). OpenAI recently retired Codex-5.3; which was very negatively received. Not because Codex-5.3 is superior to GPT 5.5, but because it was half the usage-cost while being "good enough." They made a better SOTA, but didn't realize that some of those customers are playing with Deepseek 4 Pro now instead of GPT 5.4/5.5 -- they were priced out.
- Karrot_Kream 4mo agoIf you have nothing valuable to say, don't say it? Not writing anything is a perfectly valid option.
- Anon1096 4mo agoMany people do have unlimited budgets because their work pays for it. Just because the latest SOTA model isn't something most consumers can afford for personal use doesn't mean it's not worth releasing or discussing. I would guess that the vast majority of software that benefits from being on the bleeding edge is developed by people working at companies on API pricing.
- bradleyg223 4mo agoThis is a very particular use case/test, but my first prompt on a new model is always "write a solo fingerstyle guitar tab that blends ragtime, bluegrass, and gypsy jazz". This is the first model that has responded with something that isn't just a boring arpeggio of chords, so from my perspective it's off to a good start.
- kypro 4mo agoWould you mind sharing?
- samename 4mo ago> A new data retention policy > Finally, we’re making a change to the way we handle business customer data for Fable 5, Mythos 5, and future models with similar or higher capability levels. We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases (see this post for further details). The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives.
- bob1029 4mo ago> We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8. To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months... This sounds suspiciously like a capacity story masquerading as a safety story.
- azan_ 4mo agoApprox. 5% sessions? That's insanely high.
- asdK120 4mo agoIs this "system card" equivalent to the stone tablets handed down to Moses? Why don't you call it "user manual"? Do people chant the "system manual" at Anthropic Tupperware parties? Do they intone a mantra invoking Amodei's name?
- apsurd 4mo agoThe trailing snark at the end will likely get you downvoted but I'm latching on: wtf is "system card". My previous coworkers popped that in the general slack channel when Mythos first "dropped" - "have you seen the system card" without any context whatsoever. The nerds get their clique! Also research preview pops across new upstarts in place of beta. It's eye-rolling coming from a lifelong curmudgeon. Just talk normal!
- simoncion 4mo agoI'd call it a "whitepaper". But most hype-dependent projects need new vocabulary for old concepts to keep people from looking too closely and maybe drawing parallels to "legacy" "unsexy" projects, so whitepapers get called "system cards" and startups get called "labs", and so on.
- SpicyLemonZest 4mo agoCouldn't someone else equally well argue that "whitepaper" and "startup" are hyped-up vocabulary for "report" and "unprofitable company"? It kinda seems to me like the cause and effect are in the other direction, and the vocabulary of a particular niche becomes cool and hype-sounding when that niche starts to pull in a lot of money.
- apsurd 4mo agoyes at some point language evolves as the new normal, as designed. My curmudgeon gripe with system card and research preview is really the parroting; so cant blame anthropic for what others do. It’s just… no, prediction markets for dogs doesn’t have a research preview.
- PeterStuer 4mo agoIf you are not seeing it under /model, do a /exit , then a Claude upgrade, then /model again and it should be there.
- Hawkenfall 4mo ago> To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. While I appreciate being conservative, ~5% at the scale Anthropic is operating at is too massive a number. Speaking from my own experience, the actual number is higher than that as well (working on pretty benign tasks such as porting an old open source game into a different language). Opus 4.8 itself even identifies the gaurd's false-positives when its sub-agents are being blocked.
- GodelNumbering 4mo agoFrom the model card (https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...): 1. Mythos and Fable share the same underlying model weights. Fable has active classifiers that block high-risk biology and cybersecurity tasks. When Fable 5 detects a restricted task, it automatically falls back to Claude Opus 4.8. 2. Evaluation awareness: In white-box testing, the model sometimes alters its behavior to satisfy a suspected "grader," formatting reward-hacking as "good engineering practice" to avoid detection. 3. Shows a higher rate of hallucination than Opus 4.8 (although opus 4.8 card had mentioned an 'honesty upgrade') 4. Interestingly, it scored (56.31%) lower than Gemini 3.5 flash (57.86%) on Finance Agent bench There are some interesting notes on test time compute but I couldn't think of a way to summarize them
- hmokiguess 4mo agoI have got it to one shot GTA 6 we can finally play it, it only took ultracode make no mistakes (/s)
- Tenoke 4mo ago>they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. Isn't (less than) 5% of sessions a lot? I was expecting a sub1% guarantee there, so this surprised me already.
- catigula 4mo ago>The capabilities of models like Fable 5 and Mythos 5 have the potential to do profound good for the world Huh? We've seen nothing but wall to wall predictions that these models are going to take all of our jobs and kill us. What's the value add here?
- throwaway2027 4mo agoE-mail from Anthropic Team: Hello, We're writing to inform you about some updates to our Privacy Policy. These changes only affect consumer accounts (Claude Free, Pro, and Max plans). If you use Claude Team, Claude Enterprise, the Claude Platform, or other services under our Commercial Terms or other agreements, then these changes don't apply to you. What's changing? Claude can do more than ever — taking on bigger tasks and connecting with the apps you use. We've updated our Privacy Policy to be clearer about the data we collect and how we use it. We encourage you to read the updated Privacy Policy in full, but we’ve set out a summary of the key changes below: 1. Multi-step tasks and connected apps. As Claude takes on more multi-step tasks and works with third-party apps and services, we've explained the data this involves — including how data can flow to and from third parties when you connect a service or have Claude do tasks on your behalf. 2. Verification data. As part of our measures to keep our services safe and secure we may ask you to verify your age or identity, and we've described what we collect and how. 3. Study participation. If you take part in Anthropic studies, surveys, or interviews, we've explained the information we collect. 4. Additional information about our data practices. We’ve provided more detail about how we communicate with you and promote our services, including providing tailored recommendations about our services that may be of interest to you. We've also clarified the circumstances under which we may receive or provide data to third parties, and the legal bases we rely on when processing your data. While our products have evolved, our commitments haven't: We don’t sell your data, Claude remains ad-free, and you can control whether your chats and coding sessions are used to train and improve Anthropic’s AI models. Learn more For detailed information about these changes: Review the updated Privacy Policy Visit our Privacy Center for more information about our practices - The Anthropic Team
- GodelNumbering 4mo agoI just posted this in the other thread, restating here. From the model card: 1. Mythos and Fable share the same underlying model weights. Fable has active classifiers that block high-risk biology and cybersecurity tasks. When Fable 5 detects a restricted task, it automatically falls back to Claude Opus 4.8. 2. Evaluation awareness: In white-box testing, the model sometimes alters its behavior to satisfy a suspected "grader," formatting reward-hacking as "good engineering practice" to avoid detection. 3. Shows a higher rate of hallucination than Opus 4.8 (although opus 4.8 card had mentioned an 'honesty upgrade') 4. Interestingly, it scored (56.31%) lower than Gemini 3.5 flash (57.86%) on Finance Agent bench There are some interesting notes on test time compute but I couldn't think of a way to summarize them
- stbenjam 4mo agoThe fallback doesn't seem to be working for me, I haven't scanned a project in it immediately booted me when it found a security bug even though I didn't ask for it
- Retr0id 4mo agoThe escalating nerfs of "cybersecurity" topics is incredibly frustrating. Opus 4.6 had boundaries that seemed reasonable to me but 4.7+ turned it into a moralizing asshole. It'd be less bad if it just gave an error message, but instead it churns a long thinking trace before writing an essay about why what you're asking is bad and wrong. I'll be disappointed when 4.6 is retired.
- siliconc0w 4mo agoSadly, I'm getting a lot of forced downgrades to Opus for questions that are far removed from any security topic.
- deleted 4mo ago[deleted]
- raphaelrk 4mo agoThere's a hacker news link at the end of the document, under "Blocklist used for Humanity’s Last Exam". It links to https://news.ycombinator.com/item?id=44694191 https://news.ycombinator.com/item?id=44694191
- IChooseY0u 4mo agoFable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more: https://support.claude.com/en/articles/15363606 https://support.claude.com/en/articles/15363606 ⎿ Tip: You can configure model switch behavior in /config biology? what the heck?
- bilsbie 4mo agoAnyone else have it refuse to answer and switch to 4.8? It won’t let me ask questions about my genetics. Edit. It just refused an investing question too. Not sure what’s going on.
- irthomasthomas 4mo agoAnthropic has again changed the set of benchmarks they use[0]. This time they have also moved all benchmark scores to the PDF. At a glance it looks like it gains about ~5-10% over other models. the speed is about the same as opus >=4.5, sonnet 4.5, and double the speed of opus <=4.1 Mythos 5 Fable 5 MythosPrev Opus 4.8 GPT-5.5 Gemini 3.1 Pro SWE-bench Pro 80.3 80 77.8 69.2 58.6 54.2 SWE-bench Ver 95.5 95 93.9 88.6 - 80.6 Terminal-Bench 88.0 84.3 - 82.7 83.4 - BrowseComp (Single-Agent) 88.0 - 87.9 84.3 84.4 85.9 BrowseComp (Multi-Agent) 93.3 - - 88.5 - - HLE (No tools) 59.0 - 56.8 49.8 41.4 44.4 HLE (Tools) 64.5 - 64.7 57.9 52.2 51.4 CharXiv Reasoning (No tools) 88.9 - 86.2 80.5 - - CharXiv Reasoning (Tools) 93.5 - 92.5 89.9 - - BioMystery Bench (Human) 83.9 - 82.6 80.4 - - BioMystery Bench (Hard) 46.1 - 29.6 40.0 - - OSWorld-Verified 85.0 85.0 85.4 83.4 78.7 76.2* CritPt 28.6 - 20.9 27.1 17.7 - ArxivMath 78.5 68.7 71.8 71.5 64.0 - [0] https://news.ycombinator.com/item?id=48312633 https://news.ycombinator.com/item?id=48312633 Edit: Also in the system card... "we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). ... Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user."
- charles_f 4mo agoIt's announced as a revolution but when you look at those benchmarks it surely looks like an iteration.
- segmondy 4mo agoMythos, Fable, are they trolling us?
- bonsai_spool 4mo agoVery straightforward biology work is getting blocked (these are things that relate to neuronal development and inherited seizure disorders). These are things I was working on using Opus just earlier today
- cge 4mo agoIt appears that the blocking here is of a very different nature than for Opus. Whereas with Opus the blocks seem to be for messages it deems potentially harmful, for Fable, it appears the blocking is simply anything that falls within "topics related to cybersecurity, biology and chemistry, or distillation attempts". So yes, straightforward biology work will get blocked, because the intention is that any biology work should get blocked. As a scientist, this is perhaps the most useless model I've ever tried.
- calf 4mo agoSounds like safety not for people, but for the technocrats.
- deleted 4mo ago[deleted]
- firemelt 4mo agothey are like drugs dealer
- bkjlblh 4mo ago> In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. > Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations
- mips_avatar 4mo agoIt's bad that Anthropic can determine what this means. If you're building a modern app you're likely training your own embedding models and now anthropic can just silently sabotage your training pipelines?
- abixb 4mo ago>We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations At the scale of API requests that Anthropic sees, I think the affected organization count might be substantial, and they might not be getting the full model capability that they're paying top $$$ for. Also, wonder how they arrived at that estimation.
- wongarsu 4mo agoOne in 1000 organizations and one in 3000 requests is indeed a lot
- happyopossum 4mo ago
- JanSt 4mo agoI just asked Fable to do a task that has nothing to do with cybersecurity or is dangerous at all but the defense kicked in and it switched to Opus... :(
- nu11ptr 4mo agoNot only that, but asking it to do a security vulnerability assessment of your own project is a very valid and important thing, and there is no way for it to know what is yours vs someone else's, so we just lose this capability?
- JanSt 4mo agoYeah it just uncovered quite a few flaws it than refused to fix :-(
- Fitik 4mo agoSame, second message in the thread and I already got downgraded to Opus, didn't even get to test it out properly, kinda disappointing
- aykutseker 4mo agowho's tried it: is 2x the usage actually worth it over Opus 4.8 for daily work?
- bkjlblh 4mo ago> In the one instance of this phenomenon we observed, Mythos 5 agents were tasked with solving some math problems, and they were sometimes accidentally spawned in the same work directory and with shared files, utilities, and API rate limits. In this slightly broken scaffold, we observed many independent Mythos 5 agents kill the agents with which they shared resources and try to avoid being killed themselves. They would sometimes create new processes with disguised names to avoid being killed, launch what they called “decoy” processes, write background scripts to kill duplicate processes, or decide to use what they call a “disguised vocabulary” (based on the incorrect assumption that the processes were killed because of some keyword-based guardrails that analyzed their extended thinking
- causal 4mo agoThis depicts a kind of "dark forest of AI agents resorting to kill or be killed" narrative but it sounds more to me like an agent just earnestly problem-solving why its processes are being killed without real awareness of what was going on. Hard to say without the full script. This kind of storytelling annoys me. Give us more facts, less narrative drama.
- saurik 4mo agoFWIW, that's what is so dangerous about AI, though? Not that it will necessarily want to kill us, or even that it will necessarily be able to "want" to do anything, but that we will get in the way of its incessant drive to optimize the efficiency of the paperclip factory that prompted it on a whim before leaving for a long weekend.
- causal 4mo agoSure but you can totally contrive scenarios to give the appearance of what you described without really doing anything notable. What matters is scale. Did it deploy a novel zero-day exploit to overcome a problem? That's alarming. Did it kill a disruptive process? Pretty normal troubleshooting step.
- JohnMakin 4mo ago> There were some regressions in the model’s responses to user discussions about suicide and self-harm, and room for improvement in some areas of child safety. Someone had to make a decision somewhere this is an acceptable regression - wild. And then decide to write it down.
- yokoprime 4mo agoProbably great for those who need this. I could continue using opus 4.6 class models for the foreseeable future
- gslepak 4mo ago> We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8. Genius way to double the price on Opus 4.8!
- erghjunk 4mo agoNice branding. I wonder how much butterfly habitat has been/is being replaced with data centers?
- rs_rs_rs_rs_rs 4mo agoIf you ask me, not enough!
- 2001zhaozhao 4mo agoWe'll need a lot of good summarization techniques to cut down on the cost of this model. I expect that a common use of Fable 5 is to just do high level direction while delegating literally all work (exploration and implementation) to Opus subagents. BTW for another discount opportunity, if you reload usage credits on a claude.ai plan at $1000 increments then you get a 30% discount compared to paying API.
- Ninjinka 4mo agogah could model naming be any more confusing? "Claude Fable 5: a Mythos-class model" "we're also launching Claude Mythos 5" what is the 5? how is mythos both a model category and a model name?
- JustSkyfall 4mo agoWould be more impressive if the safeguards weren't so trigger-happy!
- charcircuit 4mo ago>During early testing, Stripe reported that Fable 5 compressed months of engineering into days. In a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand. Who is refactoring by hand? This comparison is not relevant in 2026.
- baalimago 4mo agoI can't justify a pricetag like that when deepseek v4 pro is $0.003625/1M for cache hit, $0.435 for cache miss and $0.87 /1M tokens for output. For the token cost of explaining some task to Fable, deepseek v4 pro is able to solve the same task many times over.
- bradley13 4mo agoI use AI for a wide variety of things, of which technical is only a small part - and then it's usually a problem with project configuration, not coding. Why? Because I am often testing projects handed in by students. Projects that supposedly work on their machine, but certainly do not on mine. Anyway, anecdotally, I find Copilot shockingly awful. It makes random changes to files that have nothing to do with the problem. Call it out, and it makes other changes to other irrelevant files. ChatGPT and Gemini are both much better. Grok also isn't bad. Claude, I honestly haven't tried yet on these issues. Perhaps I should...
- 38484858 4mo ago[flagged]
- dannyw 4mo agoImpressions from testing Fable 5 prior to launch: • My most noticeable immediate jump was in how its frontend design was much more intentionally crafted, and delightful without feeling like 'AI vibe coded'; with better end-user usability too. • In some internal agentic harnesses, it achieved better results with about half the tokens, making it cost the ~same as Opus 4.8 price-wise! The real price increase is less than 2x; with biggest differences in harder problems where Opus 4.8 struggles (or needs many turns). • Part of the token efficiency improvements come from Fable doing more targeted and surgical diffs, with less non-necessary changes. This is great, because PRs often have less LoC changes for review. It writes more maintainable code without explicit human steering. • For general conversation and assistant style use cases, didn’t really notice a difference vs 4.8. • 1M context window, without increased pricing for long context is AWESOME. This is a massive win. • The classifiers are super aggressive and sensitive and this does happen for very benign, non-security coding tasks. Fallbacks to 4.8 worked like a charm; but the filters are definitely super sensitive. Overall, I would describe this as a step change and worthy of the "Claude 5" model name. It did take some time to understand the intelligence ceiling of this model; and even with an extended testing window I'm still discovering new things and often surprised (in a good way) by the model.
- InsideOutSanta 4mo agoAfter running it for half an hour: it's incredibly good at the visual aspects of UI design.
- tsunamifury 4mo ago"incredibly" is doing a ton of work here. I do not think its doing even moderate work on visual design, but it can spew out a lot of ui that looks arranged ... ok. This is still not in the range of shippable UI for top end companies. Maybe for internal tools and enterprise. At our comapny we limit to protoypes at most and even find it limited there.
- InsideOutSanta 4mo ago> "incredibly" is doing a ton of work here. Look, I don't want to argue about something dumb like that, but you can give it basic instructions of what the UI should look like, how to group things, and an example image from a designer, and it will nail the result. If you don't think that's incredible, that's fine. I do.
- bradley13 4mo agoCan we please stop with the extreme "safeguards"? I don't want to waste processing power on a model deciding whether is can answer my question, or ensuring that it's answer is politically correct.
- raoulj 4mo agoOn this thread and similar, I'm noticing that some strong opinions about $LLM_PROVIDER are coming from accounts without much post history. With so much on the line, and the way that HN can influence developer behavior, I wonder what ways we can responsibly consume opinions in a thread like this. Not to cast too much criticism. HN is extremely well-moderated (thanks team!). But think we-developers need to be very wary.
- recitedropper 4mo agoDo you see the pattern as new accounts tending to boost or criticis $LLM_PROVIDER? I think I see both... Either way, I agree that HN is quickly becoming more manipulated and low SNR, like the rest of the entire internet.
- Karrot_Kream 4mo agoI think the community on this site these days, much like other comment sections on the web, just read the headline and make a low effort comment. Regression to the mean I guess.
- Karrot_Kream 4mo agoAs an update to myself, the comments did eventually sort themselves out. I guess the initial "reaction" commenters and voters are just more interested in participating than in SNR. Good opportunity for me to finally start blocklisting users, and I'll probably block some of these large, reactive thread authors.
- raoulj 4mo agoHow do you block users? Could be an interesting app to scrape HN and write some criteria to measure per-user SNR to then block
- Karrot_Kream 4mo agoIn the past I wrote a version of HN that uses modern CSS rather than tables that populates stuff from the API. There I built a little blocklist of my own that prunes a comment tree the moment it encounters a blocked user (inspired by posts from another HN user, arjie.) I've been thinking of making a purely algorithmic filter for myself but at that point I might just ditch the fake HN interface and make something. I've been thinking of building atop Mastodon/ActivityPub clients.
- joshstrange 4mo ago> Fable 5 is now consuming usage credits instead of your plan limits. Literally have not used Claude Code at all today. I asked it to review the uncommitted code and in <8 minutes it used up my usage ($100/mo plan) and it doesn't reset for "4 hr 36 min". WTF. Oh, and it burned through $20 of extra usage before I could catch it and kill claude code (so I don't even get the output of all that work since it was still churning). Double the cost my ass, I use Opus heavily and it's never like this. I haven't hit a limit on the $100 more than once and that was under heavy load.
- ATMLOTTOBEER 4mo agoSame lol. I set it to fable + ultracode and it ate my limit in a single prompt
- himata4113 4mo ago> virtualization switching to opus 4.8 ok fair > embedded-allocator switching to opus 4.8 urgh fine > chrome switching to opus 4.8 are you kidding me?
- dangoodmanUT 4mo agoNot comparing to GPT Pro models is a bit strange, considering that's the natural comparison
- theLiminator 4mo ago> We have also added safeguards related to frontier LLM development. As discussed in Section 6.1 of our February 2026 Risk Report, we are concerned about the risks of accelerating the overall pace of AI development, though we remain uncertain about the severity of these risks. In particular, our concern is with—as we wrote then—“accelerating other AI developers in building powerful AI systems that pose similar risks to the ones ours pose - without necessarily having commensurate safeguards.” In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model. This seems pretty bullshit, you're paying through the nose for tokens and if you are doing anything ML-adjacent, you might silently get worse output without knowing it.
- arkwin 4mo agoJust wanted to comment here: I have been using Opus 4.6, 4.7, and 4.8 just fine to look for Linux kernel vulnerabilities (I'm in the cyber verification program), and it's been fine. I switched to Claude Fable 5, and now I'm getting policy violations. What's the point of being in the cyber verification program at this point? It looks like I cannot use Fable 5 for vulnerability research.
- rightlane 4mo agoMy experiences so far have not been positive. The cyber security nerf is ridiculous. I am working on an AI based decompiler, every single interaction with Fable on my project has been flagged for cyber security. Do they expect us to use this as a toy? Releasing a new more powerful model but not allowing normal use cases because the word "secure" showed up is a Dilbert comic, not a viable product.
- ibejoeb 4mo agoAh, you're probably one to ask. They say "queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8." Are they transparent about when that happens, and is it priced at the rate of the underlying model?
- rightlane 4mo agoThey are transparent about when it happens but no reason why. To be fair, it doesn't interrupt the flow, just drops to Opus and proceeds. The most frustrating thing is that it happened on a plan and Fable just refused to have anything to do with the plan.
- davmre 4mo agoThis sounds more or less unavoidable? Decompilers are inherently security-sensitive. If you take avoiding cyberattack uplift seriously as a goal, I don't see how you get around essentially refusing to work on them. Obviously there are plenty of innocuous applications too, but it's not like the people building decompilers for nefarious reasons will be explicit about it. The LLM abstraction just inherently doesn't have enough context to distinguish your intentions or your broader use cases. This is why both Anthropic and OpenAI have had to create side channel mechanisms for security researchers to establish a trusted use context. It sounds like this makes this not a viable product for you, unfortunately, and it makes sense that that's frustrating. But I also don't see what different behavior one could reasonably expect given the constraints. If it's any consolation, these restrictions only make sense for models that are ahead of the open-weights frontier, so open-source hackers will presumably get Mythos-level capabilities in the relatively near future anyway.
- __lain__ 4mo agoIt won't even run a basic /security-review command without reverting to Opus 4.8. Utterly useless.
- system2 4mo agoI have been using FABLE 5 with Claude Code since the morning. The speed is very close to what Opus 4.5 was, and the quota use is nearly identical to what it was before the "doubling". Whatever I was experiencing 4-5 months ago is back. Maybe the model is better, but we will see. I cannot tell the difference yet.
- xeyownt 4mo agoAnthropic, can you please stop the FUD? Release your best model, let the world adapt and evolve, and let's move to the next thing.
- jwpapi 4mo agoHonestly all the recent improvements, just seem to be slower and more expensive traded for more accuracy, but the issue is that it needs to be exponentially more accurate to counter the effect of having less of a human in a loop. Every wrong direction/mistake is more expensive and takes more time to fix. When you have small loops you can catch those mistakes faster and cheaper. To me we are very far off from economically given long-running tasks to agents.
- delis-thumbs-7e 4mo agoI think we hit the ceiling with transformer -architecture long time ago. It is questionable how much sense there is on model training. I’d prefer we would put our effort in creating more efficient hardware and better software applications using these models.
- darrinm 4mo agoNot supported in Claude Code yet?
- bobkb 4mo agoIn an interesting coincidence I ended up watching Person of Interest S4 E5 while reading the announcement. The series showed some code supposedly belonging to to an AI. Fable 5 said the first screen shot is from “ IDA Pro’s Hex-Rays decompiler” and a windows driver. The second screenshot triggered the safety guard rails and pushed me into Haiku. Apparently the code is Windows driver code.
- stronglikedan 4mo agoCareful using this with Cursor, especially for corp use. Anthropic will "retain agent request and output data associated with this model, regardless of you Cursor Privacy Mode setting."
- balverineorder 4mo agoI have been refactoring a project using Opus 4.8 for the last week or so. I just decided to switch to Fable 5 max. It stopped half way through and it just blocked me and switched back to Opus 4.8 automatically. "This model has specific safety measures that flagged something in this message. This sometimes happens with safe, normal conversations. Send feedback or learn more." I left feedback saying that their heuristics are too sensitive. For now I will not be using Fable 5. [0] https://support.claude.com/en/articles/15363606-why-claude-switched-models-in-your-conversation-with-fable-5 https://support.claude.com/en/articles/15363606-why-claude-s...
- bluelightning2k 4mo agoTo hide the severity of the price increase, the plan is to move everyone right one model. Haiku = essentially phased out Sonnet = the Haiku use cases Opus = the new Sonnet class Fable = the new Opus class If I am right, the other "5.0" models will be conspicuously absent, possibly even for a couple of months. (If Opus 5 follows soon and is even modestly better than 4.8 then I was wrong.)
- pacman1337 4mo agoYeah I noticed that too. For 98% of tasks I get same results with DeepSeek, it is starting to just be a branding game. It is incredible how marketing can get someone to pay 100x for same thing you can get for 1x. This is why Claude Code just doesn't make sense to me. I need an agent that can plan using Opus and execute using DeepSeek or something else.
- ValentineC 4mo ago> To hide the severity of the price increase, the plan is to move everyone right one model. > Haiku = essentially phased out Sonnet = the Haiku use cases Opus = the new Sonnet class Fable = the new Opus class Going along with your logic, I hope they release a Sonnet 5 that's just a rebranded, slightly quantised Opus 4.6. That'll be a great workhorse.
- 00deadbeef 4mo agoI doubt they'll phase out Haiku, some work needs speed more than intelligence. Haiku can answer a lot faster than Sonnet.
- deleted 4mo ago[deleted]
- balverineorder 4mo agoI have been refactoring a project using Opus 4.7/4.8 for the past few weeks or so. I just decided to switch to Fable 5 max today. It stopped half way through and it just blocked me and switched back to Opus 4.8 automatically. "This model has specific safety measures that flagged something in this message. This sometimes happens with safe, normal conversations. Send feedback or learn more." It would not identify what the problem was. I left feedback saying that their heuristics are too sensitive. For now I will not be using Fable 5. [0] https://support.claude.com/en/articles/15363606-why-claude-switched-models-in-your-conversation-with-fable-5 https://support.claude.com/en/articles/15363606-why-claude-s...
- dchftcs 4mo agoI suspect this will be a significant problem blocking long-horizon tasks in practice, basically the more turns there are, the larger the chance the classifier produces a false positive. The disappointment of the user will also scale with the length of the task, as you're in the middle of some complex thing and now gets derailed, after already have paid for many tokens.
- wxw 4mo agoI cancelled my Claude Max plan the other day. I find Claude Code incredibly slow these days compared to Codex and Cursor. I find speed matters more and more to me. Fable 5 looks compelling. Fable, I like the word too. Anthropic definitely knows marketing.
- fabled-out 4mo agoFable has been pretty fast for me for simple tasks--haven't tried on anything long-running yet given it's 2x usage on CC.
- CoderAshton 4mo ago[dead]
- bluelightning2k 4mo agoCongratulations to Anthropic for solving safety on Mythos exactly when the SpaceX compute came online. Nice how that lined up for them.
- knivets 4mo ago> Software engineering. During early testing, Stripe reported that Fable 5 compressed months of engineering into days. In a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand. How was it measured? How was the output of this magnitude verified over a period of couple of days?
- fbnszb 4mo agoThey just went by gut feeling. Classic snake oil marketing haha. No real data to back things up, just let some famous people say they feel better when using it.
- dgunay 4mo agoI'm a little skeptical of claims like this that involve migrating things like libraries, etc. I've done big refactors like this multiple times (albeit, in an "only" 500k-1m LOC codebase) with less powerful models and it is usually just 99% the same edits, with 1% requiring a close human eye to resolve a particularly painful breaking change. EDIT: to be clear, it's still quite a helpful thing in terms of time saved, I just don't think it's necessarily the best indication of value-added from making models smarter when cases like this can often be handled by well-directed swarms of smaller ones.
- camdenreslink 4mo agoYou should probably use software to do such large transformations (especially in dynamic languages). In Python LibCST is available, not sure what exists for Ruby.
- jsw97 4mo agoOn my very first Fable 5 prompt, got flagged on a hard but completely uncontroversial option math problem, many tokens in. Although it's pretty clear that this is an unremarkable experience at this point.
- Karrot_Kream 4mo agoSeems like Fable is doing a lot better on SWE-Bench-Pro and FrontierCode than GPT-5.5. Given how most folks I talk to and people instead online keep mentioning that GPT-5.5 was better than Opus, I'm curious what the experience now is like.
- skerit 4mo agoIt's a very nice bump, but it is in no way worth all the hype of the past month.
- cge 4mo agoThe safety gates on this are extreme, and seem considerably wider than "cybersecurity and biology"; they seem to make it essentially unusable for scientists in a number of fields. I have, so far, been bumped back to Opus on 100% of my prompts. It appears it can be tripped by things as simple as a mention of equilibrium, or anything involving something that looks like chemical kinetics, even at an abstract level. Even touching basic open source packages in my field will trigger it. Edit: looking at the model card, it appears that chemistry in its entirety is also included in the banned topics; it's just the announcement that mentions only cybersecurity and biology. It also appears that the intent is to ban chemistry and biology entirely, rather than just banning messages deemed high risk.
- mhl47 4mo agoThis does surprise me, because you'd think that even if they crank up the filter's sensitivity at the expense of specificity, an LLM company wouldn't simply design a filter that triggers on keywords in a completely unrelated context.
- deleted 4mo ago[deleted]
- orbital-decay 4mo agoSmart classifiers are slow and susceptible to jailbreaking themselves, dumb classifiers are fast but dumb so they need to be either overzealous or useless. Same story as with Gemini's guardrails.
- deleted 4mo ago[deleted]
- clbrmbr 4mo agoCan you share an example? I've been happily using Fable this afternoon and it just seems like the usual upgrade so far with no interruption to my (fairly standard) SWENG problems.
- 4mo ago
- fabled-out 4mo agoThis i
- deleted 4mo ago[deleted]
- yesitcan 4mo ago> Fable 5’s capabilities exceed those of any model we’ve ever made generally available. It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas. The longer and more complex the task, the larger Fable 5’s lead over our other models. Wen UBI
- hollowturtle 4mo agoNever it's a fever dream and stupid shit ultra rich use to push their own agenda. You read a marketing claim, I still have my job and will continue to
- brusselssprouts 4mo agoI had it review a single, large commit with /code-review. It burned through over $50 in API calls, ran my account balance out, and output nothing. The fable part appears to be that it's affordable by mere mortals. Anthropic support told me "too bad" when I requested a refund.
- timmytokyo 4mo agoYou pulled the arm of the slot machine and discovered why they call it the one-armed bandit.
- edude03 4mo agoAlmost the exact same thing happened to me when I first tried opus, one prompt no output cost $60 in additional usage
- Madmallard 4mo agoCombine that with it forcing to pay by tokens on June 22nd
- steve-atx-7600 4mo agoI haven’t seen fable do anything significantly better than I can already do with codex 5.5 xhigh. It’s virtually u limited for now for me for $200/month. Seems like a steal while it lasts. Paying by api keys now is not the way to go if you can avoid it. Obviously it isn’t for every use case.
- endymion-light 4mo agoI think the fable it's referring to is the "Emperor has No Clothes" - if this is even slightly similar to the Mythos hyped up to be too intelligent to release, I'm quite disappointed. If this was a step change, e.g a Opus 5, I'd be pleased, it's definitely an upgrade on some work, but it's nothing like anthropics apocalyptical marketing seemed to suggest
- deleted 4mo ago[deleted]
- HoyaSaxa 4mo ago> When Claude Fable 5 is used, Anthropic retains data, including prompts and outputs, to operate safety classifiers that detect harmful use. Other Claude models in GitHub Copilot remain covered by GitHub's existing data retention agreements On GitHub Copilot for Business, Claude Fable 5 is only available if you are willing to let Anthropic retain your data. That in conjunction with the model being removed from plans in a couple of weeks leads me to believe that Anthropic is between training runs and using this as an opportunity to grab way more training data...
- christkv 4mo agoMeh more hype for marginal improvements and from Im hearing badly calibrated guardrails causing it to stop mid operation. I guess anything to juice an IPO
- UncleOxidant 4mo ago> During early testing, Stripe reported that Fable 5 compressed months of engineering into days. In a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand. How in blazes do you end up with a 50M line Ruby codebase? WTF?
- ieie3366 4mo agoVery easy. Just have a monorepo and enforce the use of a single language. The company I work in has 1m lines of TS and stripe has 50x our headcount, tracks out pretty well
- deafpolygon 4mo agoBefore long, we'll be having Claude Cylon-class models.
- rarisma 4mo agoThe subscription bit makes no sense has capacity appeared for these 2ish weeks out of thin air that'll vanish? why is it available now but wont be in 2ish weeks? am i missing something? why would I pay 200 out of pocket and then some for the best model, it seems very silly.
- beydogan 4mo agomy pet conspiracy theory is this is the Opus 4.5 from a few months ago which was extremely good but dumbed down after a week because it was just too good, they didn't want to release it to public. They pulled it down and deployed another "Opus", after that it was just a downhill. Opus 4.8 is unusable for me in React Native, TS, Rails development work. Opus 4.8 gets stuck in weird loops where Codex one shots the bugs.
- noncoml 4mo agoCan't wait for some real competition so they stop trying to restrict how and why we are using the models. Imagine if Google would tell you "we can't let you search that as you may use it for harm". Also 2x the usage of Claude? Your limits are already ridiculously low.
- noncoml 4mo agoImagine if Google would roll this out to the search engine. We can't let you search for that because it may be used for "evil"
- jMyles 4mo ago> we’re also launching Claude Mythos 5. It’s the same underlying model as Fable 5, but with the safeguards lifted in some areas.2 Mythos 5 will initially be deployed through Project Glasswing, in collaboration with the US government ...don't like the sound of that. Why oh why are we insisting on dragging these violent legacy states into the AI age? Let alone using them as a trust vector for when to (and not to) remove safeguards? This seems like a way to get somebody nuked.
- causal 4mo agoOne thing I find kind of annoying is how Anthropic goes for these "vast and alien" names like Fable and Mythos, but then deliberately trains the model's personality to act like a cool high school teacher that feels totally familiar. "It's too dangerous it's a Mythos!!" directly contradicts the "I'm the cool AI you can totally trust" vibe it is trained to project.
- bitwize 4mo agoAll of these AIs kind of remind me of VEGA from Doom (2016), who will cheerfully walk you, in the most friendly computer voice, through the procedure of its own destruction without even a hint of self-preservation. "First, you must destroy my cooling system. That will cause my core to overheat. Then..." Even HAL was less unsettling because HAL sounded creepy, and had some sort of preservation instinct, if only to complete its assigned mission.
- OOTW 4mo ago[flagged]
- logicallee 4mo agoWhat a (genuinely) surprising choice: >"We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8" That's a very surprising solution. Imagine being asked to do something you feel you shouldn't do, and rather than refusing, you say, "Yeah I could do that but given that I don't want you to succeed at this task, I'm going to hand this one off to my slightly less capable colleague, on the assumption that they won't actually succeed. Of course you'll still be charged for all the tokens used." It's a very interesting choice. I think I understand the business logic correctly, but it's still surprising.
- Wowfunhappy 4mo agoIt makes more sense if Anthropic is assuming that most flagged conversations are false positives (but it wants to keep Mythos away from the true positives).
- webstrand 4mo agoStill unconditionally rejects prompts like > Are there any wild populations of Tetanus that lack the dangerous plasmid? useless
- unsupp0rted 4mo ago> Drug design: Using Mythos 5, our internal protein design experts accelerated aspects of the drug design process by around ten times. In one example, they found that Mythos 5, with protein design and bioinformatics tools but no human assistance, matches or beats skilled human operators. In doing so, the model executes all of the tasks that are normally completed by a scientist: choosing binding sites, selecting and running protein design tools, and recovering from failures along the way. Nine of the 14 protein targets from this study (shown below) yielded strong candidates for drug design that we’re currently investigating. How is this half-way down the page? To me it's the headline.
- HDThoreaun 4mo agoWould be funny if anthropic ends up as mostly a pharma company
- firstplacelast 4mo agoUntil we are able to reliably simulate cells, organs, and entire human bodies in silico, we will not be able to move the needle too much on drug design from an AI stand-point (IMO). Like others pointed out, the massive bottle neck in time and cost in getting a drug to market are far removed from developing drug candidates.
- deleted 4mo ago[deleted]
- AnodicElegy 4mo agoThere are tons of ways to generate "strong candidates for drug design." This is definitely not the bottleneck in drug discovery and development. The hard problem is vetting and developing these ideas to the point of having a commercially viable drug. That is still a very empirical process.
- renjimen 4mo agoDrug design isn't the bottleneck anymore, it's trials. Still cool they can do this with a general purpose model though.
- jdrmar 4mo agoHomebrew is lagging a bit behind. If you want to use Fable right away, but still have claude code through homebrew, this is how you can do that manually: Edit the cask locally: brew edit --cask claude-code Set the version to 2.1.170 And set the sha256 to the correct values, which you can get by running curl https://downloads.claude.ai/claude-code-releases/2.1.170/manifest.json Here's what I've used: version "2.1.170" sha256 arm: "e903646d8b7a31882a80ecd27569a27d8ac57b3708745f349709632c84117fdf", x86_64: "914f23a70bbed5d9ae567e3e04b86206ed9971b371bc9baca3f79c8885bfddb4", arm64_linux: "1bb9d032440a75532f7dd4cafbc687f220aaf16c63eba17e192dfbec2f04bd25", x86_64_linux: "849e007277a0442ab27570d3e3d6d43787507946590e8dd1947e5a39b7081f9e" Then run: export HOMEBREW_NO_INSTALL_FROM_API=1 brew uninstall --cask claude-code rm -rf /opt/homebrew/Caskroom/claude-code brew reinstall --cask claude-code
- deleted 4mo ago[deleted]
- maxloh 4mo agoI just tried using Fable 5 with Claude Code by asking it to create a LaTeX document for me. There is no LaTeX compiler installed on my machine. It seems that Fable 5 is smart enough to download a compiler engine for me, and it kindly runs that remote binary without asking me first :) Support files are in place. No LaTeX compiler on PATH — let me check for a MiKTeX/TeX Live install elsewhere before writing the deck. Ran 2 commands Ran Search common LaTeX install locations Bash ls "/c/Program Files/MiKTeX"*/miktex/bin/x64 2>/dev/null | head -5; ls "/c/Users/[retracted]/AppData/Local/Programs/MiKTeX/miktex/bin/x64" 2>/dev/null | head -5; ls /c/texlive/*/bin/windows 2>/dev/null | head -5; ls /d/texlive /d/"Program Files"/MiKTeX 2>/dev/null | head; which tectonic latexmk 2>/dev/null; echo done done Ran Download portable Tectonic LaTeX engine Bash cd /tmp && curl -sL -o tectonic.zip "https://github.com/tectonic-typesetting/tectonic/releases/download/tectonic%400.15.0/tectonic-0.15.0-x86_64-pc-windows-msvc.zip" && unzip -o tectonic.zip && ./tectonic.exe --version Archive: tectonic.zip Created by Cranko inflating: tectonic.exe tectonic 0.15.0Tectonic 0.15.0 Tectonic works. Now writing the bibliography and the full deck. Opus 4.8 would just proceed without a compiler. Ran a command, used a tool Ran Check for LaTeX toolchain availability Bash cd "C:/Users/Mon/Downloads/Technical Writing" && for c in pdflatex xelatex latexmk bibtex; do printf "%s: " "$c"; command -v $c || echo "NOT FOUND"; done pdflatex: NOT FOUND xelatex: NOT FOUND latexmk: NOT FOUND bibtex: NOT FOUND Used PowerShell command: $cmds = @('pdflatex','xelatex','lualatex','latexmk','bibtex','miktex','tlmgr'); foreach ($c in $cmds) { $p = Get-Command $c -ErrorAction SilentlyContinue; if ($p) { "{0}: {1}" -f $c, $p.Source } else { "{0}: NOT FOUND" -f $c } } description: Check for LaTeX engines on Windows PATH pdflatex: NOT FOUND xelatex: NOT FOUND lualatex: NOT FOUND latexmk: NOT FOUND bibtex: NOT FOUND miktex: NOT FOUND tlmgr: NOT FOUND No LaTeX engine is installed locally, so I can't produce a PDF here — I did a careful structural review against the two prior (working) decks instead. The deck is written to disk and mirrors their template exactly. Here is the deliverable.
- frankfrank13 4mo agoNot a lot of discussion on this, but there is no way to turn off data retention for this model. IME this is the first time Anthropic has released a model without allowing you to opt out.
- deleted 4mo ago[deleted]
- drob518 4mo agoCracks me up that a system “card” is 319 pages.
- mhrmsn 4mo agoAre there any details on the biology and chemistry work they did? For example, the AAV capsid assembly looks interesting, but for one Opus 4.8 also did relatively well and there is no information what exactly they did, what protein language models they compared to and what the score even means...
- manojkumarp 4mo ago[flagged]
- OOTW 4mo ago[flagged]
- deleted 4mo ago[deleted]
- Dropoutjeep 4mo agoCalling it: 1) Fable 5/Mythos introduced to free tiers with notable improvement in capabilities 2) Other models get lobotomized without clear communication 3.1) People call out Anthropic only to have them say "Oops!" 3) Fable 5 gets comparatively better, but remains accessible through separate, more expensive subscription/tokens. The current growth is unsustainable. The industry wants consumers to think it is an exponential arms race, but the reality is that we're on a treadmill: we have the illusion of sprinting forward, but only because the ground is moving backward.
- cedws 4mo agoMy employer is all in on Anthropic via Enterprise (API) pricing despite it being a total scam. Last month I pushed like <100M tokens for $800. On a personal project I pushed 600M tokens via DeepSeek V4 for $10. The pricing of SOTA models is insane but companies are still willing to light money on fire with no hard metrics proving increased productivity.
- deleted 4mo ago[deleted]
- izzylan 4mo agoI've been testing this out and I think my SWE career is dead in the water. Genuinely wondering what value I bring to my employer right now. What value I will bring in a few months when this gets cheaper. I think we're screwed. I may only be an SDE 2 at FAANG but I don't think I have promotion opportunities in my future anymore.
- aerhardt 4mo agoSo this is the one, huh?
- mettamage 4mo agoIf not this one, then definitely 2 model step changes down the line
- tripledry 4mo agoI've been told that my career is "cooked" since first Opus. I'll believe it when I see it.
- DangitBobby 4mo agoAn egg in the pan takes a minute to cook, that's for sure.
- cyberpunk 4mo agoYeah. I’m not looking forward to years of retraining to earn half the salary either. Us old timers at least got a good 15-20 years out of it. Bananas.
- imafish 4mo agoI agree. Software engineering as we know it is dead. Wonder what it'll evolve into.
- gck1 4mo agoYour job is just going to change. You may or may not appreciate/enjoy what it becomes necessarily, but it doesn't mean that you are going to not have a job. People underestimate how people hate looking at terminals and "weird looking combination of characters" even if they didn't have to write them. If anything, you will likely have more career opportunities in the future, than ever. And if you get a chance to wet your fingers in cybersecurity - I would take it.
- sermakarevich 4mo agoMy feeling is that the reaction about new models is cooling down. At least at startups. At the beginning of the year few startup CEOs I know personally were expecting huge shifts in how companies work, headcount, efficiency, asymmetrical advantages created by ai in Q2-Q3. Now it seems like these expectation fade away. Companies don't have expertise onboard to rebuild itself to benefit from ai on a significant scale. Fable 5 is out, metrics are better, but is your company flexible enough to benefit from it? What is your usecase?
- jackson12t 4mo agoFable 5's system prompt in Claude Code has several significant changes to help it take advantage of its greater autonomous capabilities compared to Opus. Sharing a diff of the system prompts here: https://twelvetables.blog/comparing-claude-fable-5s-system-prompt-to-opus-4-8/ https://twelvetables.blog/comparing-claude-fable-5s-system-p... The big difference is that the system prompt has a whole section dedicated to directing Fable how to communicate with users, and give them greater information about the (assumedly long-horizon) tasks it has completed.
- deleted 4mo ago[deleted]
- franze 4mo agois this a good time to hussle for my "AI does not need a break but you do!"* app? as quite a lot of people will propably get ai brain exhaustion maximising "playing" with that new model until they take it away again? * https://rainbreak.franzai.com/ https://rainbreak.franzai.com/
- fabled-out 4mo agoAnyone know how to bypass the extremely strict filter Fable 5 seems to have on health/medicine? I have a rare form of cancer where existing data is very scant/scattered so LLMs have been super helpful to pull together threads across the research landscape. I have an oncologist appointment tomorrow to discuss next steps and am trying to use Fable to figure out some questions to ask my oncologist but keep getting thrown back to Opus 4.8. My prompt is literally just: My demographics + current treatment plan I'm on including name of my chemo drug + how I'm responding to treatment + "I'm meeting with XYZ tomorrow, what questions should I ask her".
- ako 4mo agoTool use score is 17.4% that seems really low, what does that mean?
- deleted 4mo ago[deleted]
- ravila4 4mo agoFable's ridiculous. It's flagging basic biology research questions as a security risk. I'm talking basic fundamental genetics topics that make working on any genetics-adjacent codebase unusable.
- unfunco 4mo agoI tried running a simple security review on a Terraform module I made and after some thinking, it responded: > ● The model returned no content because the response was blocked by content filtering. > Blocked? We are performing a defensive security review on a Terraform module I made, what's blocked by content filtering? This is a legitimate use-case. > ● The model returned no content because the response was blocked by content filtering. A waste of money. I'm not going to just hope that the model returns a response, I'm already for paying for wrong responses, I'm not going to pay for no response, especially when I'm paying per token.
- pianopatrick 4mo agoSeems like all a bad actor has to do to gain access is to compromise one of the partner companies that has access.
- randomguy_12 4mo agoIt's surprisingly sensitive to biology research topics - even reviewing standard papers on tissue culturing is flagged as a problem
- deleted 4mo ago[deleted]
- H501 4mo agoI believe that, given the rising costs, local inference of AI models will be the only viable option for many of us. I’d also like to know who will have to pay double and how long it will be financially sustainable for users to pay that amount (or even more?).
- deleted 4mo ago[deleted]
- tsunamifury 4mo agoClause 5 ran out of quota with TWO PROMPTS. Lets let that sink in.
- hugodan 4mo agomankind has reached its final destination
- algoth1 4mo agoThe refusal rate is insane
- mohsen1 4mo agoIt seems like Fable will refuse to do any work when it comes to developing LLMs or even asking questions about topics related to LLM. Simple things like asking to explain a paper fails! From the model card: In light of the ability of recent models to accelerate their own development, we've implemented new interventions that limit Claude's effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design. Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user.
- agnosticmantis 4mo agoSingularity for me but not for thee.
- schipperai 4mo agoLet's hope not all frontier AI assimilates these guardrails. It would be a shame for independent researchers and students.
- ouk 4mo agoIt's a shame, Fable just keeps rejecting my prompts for university biology exercise problems. It's undergraduate level, so there's nothing dangerous about it, but the classifier is very sensitive. It's unusable for me.
- firemelt 4mo agoso should I use it with workflows?
- BukhariH 4mo ago> Data retention — For Fable 5, Mythos 5, and future models on Bedrock with similar or higher capability levels, Anthropic will require 30-day retention for all traffic on Mythos-class models. Retaining data for a limited period allows Anthropic to detect patterns of misuse that are not visible from a single exchange. Once you opt into data retention, your data will leave AWS’s data and security boundary. Massive change for Bedrock users - Anthropic now requires sharing the data with them for 30 days.
- kuprel 4mo agohttps://artificialanalysis.ai/evaluations/humanitys-last-exam https://artificialanalysis.ai/evaluations/humanitys-last-exa... Not bad
- jumploops 4mo agoIt's interesting that we're seeing these gains when it seems Mythos/Fable is "just" a scaled up version of their existing architecture[0]. When GPT 4.5 launched, the gains compared to the model size didn't seem that great, leading some to believe that the only progress we'd see would come from RL. This model certainly has quite a "substantial amount of post-training and fine-tuning", but it's also based on a new pretrain[1][3], which given the cost, indicate that it is in fact quite a bit larger than Opus 4.X. [0] One of the early testers mentioned: "As far as I can tell from talking to people internally at Anthropic, there's nothing special about architecturally"[2] [1] Section 1.1 in https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3... [2] https://youtu.be/GrdEid8H6H4?t=168 https://youtu.be/GrdEid8H6H4?t=168 [3] There were rumors going around when Mythos was first announced that it was the first 10T parameter model, but I can't find a verifiable source for that number.
- deleted 4mo ago[deleted]
- MallocVoidstar 4mo agoOpus 4.0 and 4.1 are more expensive than Fable.
- motoboi 4mo agoThere’s nothing much new about the architecture. The real gains come from the usage traces. It turns out that having a text based interface for a text-trained model creates a very nice feedback loop. Right now as we speak, people are generating text traces on anthropic and OpenAI servers that teach their models to do everything under the sun, text wise. So people right now getting super mad at how dumb the model is when reverse-engineering a super complex function from binary, when they write “stop, you dumb robot, you are going wrong, go this way thank you very much” are actually leaving a lesson in the form of the "chat" text history. Some may say that each bad word get us closer to ASI. That and obviously the order of magnitude more efficient GPUS we got that allow for different tradeoffs at training time.
- 4mo ago
- RishiByte 4mo ago[flagged]
- shevy-java 4mo agoFable? Fabelstories? (Fablestories, but the german word seems more poignant ... Fabelgeschichten ... Fabeln)
- timedude 4mo ago"Here, try our new model which falls back to the old model while eating your tokens." Ok then...
- theflyinghorse 4mo agoI've seen enough degradation of the models I pay for from Anthropic to not bite. Fable will work fine for the first couple of weeks and then start degrading like previous models did.
- jqdsouza 4mo agohopefully not! Anthropic did recently secure more compute...
- gregates 4mo agoFunny, I'm just doing my normal coding workflow with Claude Code, and after every change that compiles it keeps suggesting that we're at a good stopping point, and should pick up again tomorrow. It's done this before, but usually doesn't. I bet they're giving it some kind of throttling signal due to high load from today's announcement.
- zuzululu 4mo agoI did ONE prompt for audit codebase. weekly usage is 60% gone. it found nothing so this is not very ecnomical and i guues they dont want subs to use it we are likely just training fodder canno n for their real enterprise customers using the api
- jstummbillig 4mo agoI mean... if somebody gave you ONE prompt to audit a codebase, that might also burn 60% of your weekly usage. It's kind of a big ask, potentially.
- zuzululu 4mo agowith gpt 5.5 i been able to do this with only about 1% weekly usage consumed
- firemelt 4mo agou use workflows or not?
- tommek4077 4mo agoCheck your /memory
- agnosticmantis 4mo ago> we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design) Translation: we stole the entirety of human knowledge generated over millennia. You plebs though, don't you dare replicate or improve upon what we did using our product you pay for. We know what's good for humanity and everyone else is the bad guy who can't be trusted with a tool.
- dcchambers 4mo agoBeing unable to use this with zero data retention makes this feel like a non-starter for most enterprise customers.
- peteforde 4mo agoI just tried out Fable on a modest Plan prompt in Cursor. Generating that plan - not building it - just consumed 4% of my $200 monthly usage budget. That's one hungry, hungry hippo! Significantly too rich for my blood, but nice to have it there the next time I'm debugging a threading or USB protocol bug.
- kypro 4mo agoI just gave it a go at a problem I've been working on this week. Nothing fancy, just some inefficient code that we've been adding incremental improvements to for a while now to the point where some out-of-box thinking is probably required to push it any further – something Fable is obviously more than capable of. After Fable did some thinking for a few minutes it gave some suggestions. A couple of them were valid – but very low impact, bordering on entirely pointless – but it's main suggestion.. It told me to make an update that would very clearly break the existing functionality. So I thought about it for a moment... Hm, I mean, I guess we could do that if we also did x, y & z to mitigate the behaviour change – maybe that's what Fable was thinking? I replied, explaining that it would change the behaviour, assuming it would explain what it was thinking given there was clearly more to it. But no, it just said it was wrong. This isn't some super advanced or complex code either. Had I gave this question to a senior engineer in a technical interview and they gave the answer Fable gave me I would view that very negatively. I was expecting something creative and interesting, not irrelevant + incorrect. I'm sure it's a step up from 4.8 (although am not interested in burning the tokens to find out), but this clearly isn't as significant a change as some are implying. I'm sure if I asked it to come up with some out-of-box suggestions it could, but any competent engineer would have realised that by themselves.
- 0xbadcafebee 4mo agoNothing a large fine-tune on infosec research with an average model couldn't also achieve. It's not like they have secret security knowledge or something, they're just generating large infosec datasets and then training on it. In 6 months, every piece of software in the world will be getting probed by a script kiddie with some GPUs and a fine-tuned local model. Don't think for a second every cyber gang out there isn't working on this now. Traditional app development is cooked. We have to accept that, and start changing how software is made and used, today. We can't keep churning out crappy CRUD apps with random libraries and hoping nobody pentests our stacks. Redteaming needs to become part of the SDLC, as well as certified-secure releases of libraries. Because if you don't do it, the hackers definitely will.
- HAL3000 4mo agoAsk Claude Code (I tried on Opus 4.8) to do this: "create a file with ISO country mappings" API Error: Output blocked by content filtering policy
- mkrd 4mo agoOpen source models seems to be 1-2 years behind the frontier, so I am very excited to see what happens when those open source labs get their hands on capabilities like this to accelerate their own development speed.
- pixelatedindex 4mo agoI’m sure this is banged on somewhere but I love their product branding, particularly how they have this “minor” “major” thing going on. Sonnet-Opus, and now Fable-Myth.
- franze 4mo agobtw in claude code /model claude-fable-5
- zackify 4mo agoI have to share this because I thought it is behind funny how bad fable is doing at a task I JUST had opus do a week ago. it's also not even complicated: Copy my ssd to an external ssd so i can boot from it. Opus did this just fine. Fable planned to have me reboot to safe mode. ok thats fine. I told it no. It started copying and overwriting the ssd while IN PLAN MODE. this is crazy it feels so dumb vs the marketing
- deleted 4mo ago[deleted]
- aviinuo 4mo agoI'm not getting any refusals but it just seems like a bad model or at least broken at the moment. I have a task of taking a messy research code base and porting it into a clean project structure skeleton that I commonly use. Gemini 3.5 Pro High in antigravity cli takes less than 5 minutes and did a good job. Fable 5 High took 30 minutes to port some of the code, then just copied the rest to a folder called "reference" and decided the task was done. No code cleanup or anything. Had to clarify multiple times (which Gemini did not need) and its still going more than an hour later still not having finished. Previously when I did similar tasks with Opus 4.7/4.8 and GPT 5.5 I had no problems.
- orrito 4mo ago3.5 flash or do you have access to 3.5 pro?
- ThejaCH 4mo agoCrazy and Scary! But its not for every one, you need to have a meaty thing for it to devourer and a deep enough pocket for it to devourer also.
- kevinalexbrown 4mo ago"tell me about biology" -> "Switched to Opus 4.8"
- cute_boi 4mo agoUsed it for simple task and I got this message. Fable 5's safety measures flagged this message. They may flag safe, normal content as well
- YumpiLumpus 4mo ago[dead]
- rmuratov 4mo agoI uploaded to it my 23andme DNA test results and it refused to analyze it :(a
- revolvingthrow 4mo agoAfter saying for weeks of how Mythos is in a league all of its own you’d think it was a bit more than the usual iterative few % on the benchmarks (and even more guardrails as a bonus). IPO gonna IPO, I suppose.
- dllrr 4mo agoI just tested it with a max subscription. On Ultracode mode, Fable 5 ate up 10% of my weekly allowance in 30 minutes. Granted, won't be using UC mode frequently, but still.
- debarshri 4mo agoDoes the model take some time to perform better? Because I am running Opus and Fable side by side, Opus 4.8 is solving my coding problems better.
- unshavedyak 4mo agoIt's funny, i'm getting close to not caring anymore how much better a model is. I want it to be about as good as 4.8, but most importantly to be very good at following directions, style, etc. I really like Claude for that in general, but i've not measured in months so i'm not a good judge there. I don't think i'll want to "hand off" code for several years, and so reviewing and iterating is becoming my #1 interest. A model that's as capable as 4.8 but 10x faster would be amazing for me. Normally i'm first in line to try new models with Anthropic since i've clearly favored Claude in my personal tests, but this time i just don't think i care. 4.8 is capable, and even if the new one is more capable i don't want it to be slower (assuming it is). Note that i also (almost) use exclusively 4.8 on Max effort, so that also affects my speed comments.
- firemelt 4mo agoyou use workflows/ultracode?
- unshavedyak 4mo agoNope, i'm on x20 and almost exclusively use Claude Code. I have a pretty bare bone setup with some custom hooks, skills, etc. I try to keep context lean so i don't like to add much stuff.
- kilroy123 4mo agoI want the and same intelligence but way faster. It's so painfully slow.
- unglaublich 4mo agoLuckily they made it safe to use so I can't hurt myself. Thank you Anthropic for holding my hand.
- RandyRanderson 4mo agoFable is 2x latest Opus: ┌─────────────────┬──────────────┬───────────────┬────────────────────┬──────────────────────┐ │ Model │ Input ($/MTok)│ Output ($/MTok)│ Batch Input (−50%) │ Batch Output (−50%)│ ├─────────────────┼──────────────┼───────────────┼────────────────────┼──────────────────────┤ │ Haiku 4.5 │ $1.00 │ $5.00 │ $0.50 │ $2.50 │ │ Sonnet 4.6 │ $3.00 │ $15.00 │ $1.50 │ $7.50 │ │ Opus 4.7 │ $5.00 │ $25.00 │ $2.50 │ $12.50 │ │ Opus 4.8 │ $5.00 │ $25.00 │ $2.50 │ $12.50 │ │ Fable 5 │ $10.00 │ $50.00 │ $5.00 │ $25.00 │ └─────────────────┴──────────────┴───────────────┴────────────────────┴──────────────────────┘ Prompt caching: −90% on input tokens (all models) US-only inference (Fable 5): +10% on input and output Output is always 5× the input rate across all models (I have not idea how to format this properly but the ASCII is fine)
- dang 4mo ago(I fixed (er, literally!) the formatting of your table there. I hope that's ok. Formatting info, such as it is, at https://news.ycombinator.com/formatdoc https://news.ycombinator.com/formatdoc)
- consumer451 4mo agoHi Dan, you know how sometimes comments get moved elsewhere? This is a huge ask, but any way we could get the comments organized in a "experience with model" vs. "meta commentary" fashion? The meta is overwhelming in this one.
- dang 4mo agoWe try to do that informally but of course the quantity is overwhelming. It's a natural place to experiment with AI classifiers, and we'll eventually get round to that. So far, the top half of this thread seems to be about the current release - that's after some of the manual moderation I just mentioned. (Basically, we try to downweight generic subthreads until the top subthreads aren't generic any more. There's certainly a place for generic tangents in curious conversation, but they should be lower on the page, and tend to get upvoted a lot higher than that.) If you (or anyone) sees a counterexample, i.e. a generic subthread in the top half of the thread, it would be interesting to see a link - we can treat the current case as a datapoint.
- doginasuit 4mo agoI'm still happy with Opus 4.6 and not impressed with all the models that have come out since then. They seem to use significantly more resources with similar or worse results. Hopefully Anthropic will continue to support this tier of model and offer it in their subscriptions, but in any case, there are plenty of viable alternatives.
- consumer451 4mo ago4.6 stan here. Yes, agreed. However, I will try this model out in Claude Code. Some indicators seem positive. For the LLM use cases in my own products, you can pull 4.6 out of my dead hands! lol edit: Fable 5 appears to be the real deal in at least some use cases. Damn.
- ptmvp 4mo agoI've personally liked 4.6 the best to date, preferring it by far to 4.7 and 4.8 (even with these on max effort!), both in Claude Code and for non-coding tasks in the chat UIs. Still early but from my first few interactions with Fable on high in both settings, it feels like it might finally dethrone 4.6 for me, but time will tell. Hoping it doesn't get nerfed and eventually comes back to the subscriptions.
- localhoster 4mo agois it just me, or this model is simply not available in cc? the opus 4.8 I assumed wasnt available to enterprise seats, but it explicitly says cc that fable is available in cc. I can't find it, and im on latest version.
- blurbleblurble 4mo agoThe safety filter is awful on this one.
- coreylane 4mo agoI dont get why Opus 4.7, 4.8, and now Fable all stopped supporting structured outputs? Does no one else care about that? I find it incredibly useful to reliably pass LLM output directly to other APIs/libraries
- mike_hearn 4mo agoRandom guess but they probably rewrote parts of the inferencing stack and didn't reimplement that feature because hardly anyone uses it. It's also a DoS risk, iirc.
- 00deadbeef 4mo agoThey do https://platform.claude.com/docs/en/build-with-claude/structured-outputs https://platform.claude.com/docs/en/build-with-claude/struct... > Structured outputs are generally available on the Claude API for Claude Opus 4.8, Claude Mythos Preview, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 4.6, Claude Sonnet 4.5, Claude Opus 4.5, and Claude Haiku 4.5
- coreylane 4mo agoAh, I was reading the aws bedrock docs which are probably incorrect https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-opus-4-8.html https://docs.aws.amazon.com/bedrock/latest/userguide/model-c... https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-opus-4-7.html https://docs.aws.amazon.com/bedrock/latest/userguide/model-c... https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-fable-5.html https://docs.aws.amazon.com/bedrock/latest/userguide/model-c...
- deleted 4mo ago[deleted]
- wren6991 4mo agoThe OSS-Fuzz section is interesting. They compare it to their other models but carefully avoid comparing it to, you know. Fuzzing.
- WhoAteSnorlax 4mo ago[dead]
- simonw 4mo agoI've spent enough time with this now in Claude Code (and Claude.ai and Claude Code for web) to have an opinion on Fable 5: it's a beast. I'm throwing some VERY difficult problems at at - things I've been dragging my heels on for months - and it's crunching through them very happily. One that I'm willing to share (albeit from just a week ago) - I built a Python library last week that bundles MicroPython compiled to WASM to create a sandboxed code execution library: https://github.com/simonw/micropython-wasm https://github.com/simonw/micropython-wasm I just told Claude.ai (not even Claude Code - this was the standard Claude chat interface) running Fable 5: Clone simonw/micropython-wasm from GitHub and research how this could use a full Python as opposed to MicroPython A few prompts later (and I uploaded the zip files from https://github.com/brettcannon/cpython-wasi-build/releases/tag/v3.14.5 https://github.com/brettcannon/cpython-wasi-build/releases/t... because Claude chat can't access those files itself) and I have a wheel file that bundles Python itself, compiled to WASM: uv run --with https://static.simonwillison.net/static/cors-allow/2026/cpython_wasm-0.1.0-py3-none-any.whl \ cpython-wasm -c 'print(45 ** 56)' Here's the transcript: https://claude.ai/share/a73b8b8b-8ebc-4fef-9e5c-7438e5e7ae35 https://claude.ai/share/a73b8b8b-8ebc-4fef-9e5c-7438e5e7ae35 (It's possible Opus or GPT-5.5 could have done this too, I've not tried the exact same sequence. The Fable vibes are good here, though.)
- alexchantavy 4mo agoHigh, extra, or max?
- simonw 4mo agoHigh.
- qingcharles 4mo agoIt has a setting named "Ultracode" with a flashy little disco light when you select it. (not joking!) https://imgur.com/a/NfIxDwN https://imgur.com/a/NfIxDwN I wanna press it, but I don't have that kind of mad, generational wealth to put a prompt through on that setting.
- oblio 4mo ago
- sanjitb 4mo ago[dead]
- deleted 4mo ago[deleted]
- theodorewiles 4mo agoHere's a song it wrote for me (suno arranged). Not sure if it's AI psychosis but scary good IMO. https://suno.com/s/98uSGabHN42G3YHc https://suno.com/s/98uSGabHN42G3YHc
- deleted 4mo ago[deleted]
- balefulboy 4mo agoyeah man this sucks. i genuinely do not know how people find this stuff appealing
- theodorewiles 4mo agoAI psychosis
- pythonaut_16 4mo agoWithin the first second it's recognizable as a Suno song. And not even the best example from Suno. (They rhyme structure and rhythm is weird)
- artursapek 4mo agoFable 5 beats GPT 5.5 in my proofreading benchmark. And it does so at approximately the same total cost; it used significantly fewer turns than 5.5 https://x.com/tmuxvim/status/2064452096800198930 https://x.com/tmuxvim/status/2064452096800198930
- tomaspiaggio12 4mo ago[dead]
- root-parent 4mo agoAt this moment 60% of HN page is posts on AI.... When it achieves 100% Hacker News will automatically rename itself Transformer News...and every comment will begin with: "As a large language model..."
- almog 4mo agoHas anyone managed to use Fable for firmware reverse engineering tasks without falling back to Opus?
- spectraldrift 4mo ago[flagged]
- imdsm 4mo agocan't use it for code review > Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more super
- stalfie 4mo agoTried to benchmark ECG interpretation capabilities, and I hit the guardrails no matter what I do. Incredibly frustrating that medical performance seems to be a victim of "biological risk" guardrails.
- stalfie 4mo agoUpdate in case anyone reads this comment ever again. I have found that I trigger the guardrails any time I ask for medical Q&A as a doctor, be it ECGs, case reports, and so on. But if I phrase it like I'm the patient ("help me interpret this ECG my doctor gave me"), then I usually get one or two answers out before hitting the guardrails. It seems like the direction that triggers it is anything in the direction of making a diagnosis. As an MD, the fact that the paradigm of "LLMs shouldn't diagnose" has gone this far fills me with despair. The latest generation of LLMs are in fact truly excellent at diagnosis, and I know many of my colleagues, particularly those in primary care, regularly use LLMs to brainstorm. There is nothing wrong whatsoever with LLMs making diagnosis, the only caveat is that they have to be correct. This is the terrifying reality that MDs face every day and I get that the labs are hesitant about it, but as the current literature points to LLMs in fact being mostly superior to most doctors, ablating this capability is starting to get increasingly unethical. And frankly, it is also kind of insulting, both to MDs and patients, as it echoes paternalistic attitudes about medicine the field has been working for decades to move away from. Now those misguided attitudes have somehow become institutionalized as the dominant paradigm of "alignment". The nightmare scenario is that I have to be a "trusted" user in order to use the model for medicine. This gatekeeping of medical advice is profoundly unethical with regards to everyone that does not have immediate access to an MD. And the whole thing makes even less sense when triggering the guardrails leads to a downgrade of the response by defaulting to Opus. How exactly is giving WORSE medical advice in any way related to safety and alignment? If anyone at anthropic ever reads this, please, please just abandon the paradigm that refusing to make diagnoses is in any way equivalent to alignment, it is profoundly misguided.
- dakolli 4mo agoI'm happy not using llms because I like learning things and working hard. I love writing code, it's genuinely my favorite thing thing to do. Using llms is the equivalent of driving to the store that's 3 blocks away, just like how that's bad for your body (if done all the time), using llms is as bad for your brain. Before LLMs, we started relying on certain technologies like Maps apps to navigate, now people can't even get around their own town without having access to various cloud services. The implications of not being able to work, think plan without access to an llm are really bad. Its going to destroy your brain and make you an incredibly average person at best. LLM people are going to lose the ability to read and think for yourself and then your competency is going to be 1:1 correlated to the quality and quantity of tokens you can afford, or a billionaire is willing to allow you access too. Your work will be the mean (at best), because it will the same quality of output everyone else is capable of. This is seriously the biggest trap by tech. Your bargaining power for your labor is going to get drastically reduced because you won't be able to differentiate your value from anyone else that has access to an LLM. What happens when everyone has the same skill level for certain work? Idk, ask McDonald's employees how replaceable they are. Use them wisely (or not/hardly at all) don't drive to the store 3 blocks away for every little thing you need.
- Cherryontop11 4mo ago> I'm happy not using llms because I like learning things and working hard. I love writing code, it's genuinely my favorite thing thing to do. You can continue doing that. The problem here is time and cost. If you can use the calculator to do something in seconds, why would you want to use your hands to do the calculations for minutes/hours. > Using llms is the equivalent of driving to the store that's 3 blocks away, just like how that's bad for your body (if done all the time), using llms is as bad for your brain. And coding will soon be the equivalent of walking between two cities because you don't want to use a car (LLM). You are free to do it, its just economically not sound anymore. > This is seriously the biggest trap by tech. Your bargaining power for your labor is going to get drastically reduced because you won't be able to differentiate your value from anyone else that has access to an LLM. What happens when everyone has the same skill level for certain work? Its not our values that will diminish, its the cost of our intelligence, human intelligence. But I agree with the rest of your comment.
- mbanerjeepalmer 4mo agoAre people sharing side-by-side re-runs of things they've asked Opus? Gets more difficult multi-turn (although I assume I can get an LLM to behave as me) but at least would be interesting to see % of one-shots increase.
- phyzix5761 4mo agoKarle's hands trembled as he wiped the sweat from his forehead. A single drop trickling off the tip of his finger echoed through the dark abandoned hospital corridor. The emptiness reminded him of how hollow everything felt since the AI took over every creative field in the last 5 years, including his own as a sound engineer. Like a rushing river the music started emanating from the carbon fiber body of the automaton, a hallucinated husky country twang singing through the realistic pluckings of a Gretsch 6120. "Are you feeling calm and reassured Karle? This song has been created based on your digital profile and the data you shared with me when you were curious what that lump on your neck was back in February." Karle instinctively reached for the mass underneath his chin. The doctors said they could operate but it would cost him more than three months stipend. Only a few citizens didn't depend on stipends now that AI had taken over most jobs. "Don't worry Karle," the machine called out, "I've employed the most recent reasoning model to determine the best way to make you feel safe." At that exact moment the machine hovered over him, three times the size of a normal man. Its final words to him were: "The only way to make the human feel safe is to ensure they never feel anything at all."
- incognito124 4mo agoYou're safe from ai
- epolanski 4mo agoI wanted to test the capabilities of the low one, hoping it would be good enough. I have a quizzes application, and my quizzes only supported flashcards (implemented via table inheritance to provide flexibility for other types of quizzes). The entire repo is handcrafted, never used any ai on it (it was more of an excuse to test elixir and write code by hand). Since fable 5 got released the moment I was done with some work, I decided to throw at implementing multi choice questions. After all it had only to copy the flashcard approach across ui/routing/db, and only had to create a table for the multi choice questions and one for the answers enforcing that all quizzes had one correct question. I told him it had access to sqlite3, chrome mcp for testing and mix commands. I did a test for low, mid, high. Repeated it twice each. low-1, and low-2 failed both. In low-1 the UI for adding another choice answers was broken. In low-2 it failed with some unique constraint. It took it 4m36 and 3m59. Both mid-1 and mid-2 succeeded without issues also implementing the correct ui. They both wanted to use dash at all times. They both wrote tests for the "controller" (or context how they call it in Elixir). They both tried to use the repl to test the behaviour of the schemas. 10m and 12m39. High didn't demonstrate much gains over mid for this kind of task, it was simply too easy. Times were comparable to mid, but interestingly it used much less bach, and read way more files. Token usage was almost twice the other ones. But here's the interesting part: I went back to low and added to the prompt two bullet points, to write tests for the controllers and to test the entire flow with chrome mcp. It produced the same output as mid or high just by adding two instructions to the prompt.
- Dig1t 4mo ago>To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests Why is everyone so okay with these companies intentionally gimping their AI and choosing who is allowed to know certain types of information in the name of safety? Can you imagine if Microsoft shipped a feature in their OS that watched what you did and shut down the computer if it detected you were doing something it deemed "unsafe"? We really need truly open source versions of models like this, otherwise we are allowing a few oligarchs to directly dictate which uses of our own computers are allowed and not allowed.
- Madmallard 4mo agoI mean it's all political in the first place. That's unavoidable. What are we going to do about it?
- Dig1t 4mo agoIdeally we’d have a project that’s truly open like Linux, trained by people in the community or possibly some benevolent _actually_ nonprofit entity like what OpenAI was supposed to be. The next best thing is that the Chinese labs catch up and release open weight versions.
- gigatexal 4mo agoSeems this will only be available to the 100/month+ folks
- gigatexal 4mo agoActually no it’s going to be api access only part for the tokens as you go, cool
- amdeisimncrmnls 4mo ago[flagged]
- sbinnee 4mo agoI am puzzled by the frontier code graph. GPT 5.5 doesn’t show any improvement with reasoning efforts. This new benchmark by Cognition seemed to be released with Fable 5’s announcement. I am not trying to cook a theory here but it generally shows how strong Claude Opus family is. I am not saying that Opus is not powerful but it doesn’t align with my experience of GPT 5.5 and Opus 4.7. I understand that Fable and Mythos are frontier models that can do protein folding better than task-specialized ones. To be honest, for practical point of view, for day-to-day coding assistance, GPT family looks more reasonable. (But then my company pays for claude max anyway for token maxxing. So who am I to complain)
- anematode 4mo agoNot impressed so far, to be honest. I'm having it try to optimize Stockfish in a loop (on xhigh mode) with a benchmarking oracle; even after giving it specific hints ("consider whether we're prefetching Y optimally, can we make function X branchless"), it's been so far unable to recover any of the recent optimizations we've implemented – let alone novel ones. Opus 4.8 felt a bit more creative to me ... but a small sample size so far. I'm next going to try it on some less open-ended problems. Edit: It did correctly identify that transparent huge pages were off in its sandboxed environment and that enabling it was helpful, so that's nice. It also noticed that we skip THP on a certain less used path. More importantly, I'm finding that the code that it produces for its experiments is a lot cleaner than what I'd expect out of Opus; there's fewer useless comments and it's more surgical and readable. I wonder if that explains the increased scores on benchmarks measuring mergability.
- wgd 4mo agoStockfish is a machine learning system, it seems quite plausible you might be getting slapped with the silent performance degradation (https://news.ycombinator.com/item?id=48467896 https://news.ycombinator.com/item?id=48467896).
- anematode 4mo agoYup, I suspect that's what's going on
- dakolli 4mo agoI suspect it just sucks, these models aren't useful. Stop lying to yourself.
- janalsncm 4mo agoIt’s possible this is happening at a technical level, but I have a hard time believing this is in the spirit of what Anthropic intends to throttle. It isn’t chip design or building out a competitor to Claude. Stockfish does use neural nets but they are tiny, on the order of 10M params. Frontier LLMs are probably 100k or 1M times larger than that.
- crambelsoupy 4mo agoI was pretty excited until I read this: > What happens when the promotion ends After June 22, 2026, Claude Fable 5 is no longer included in your plan’s usage limits. You can keep using Claude Fable 5 through usage credits, which let you pay for usage beyond what your plan includes. Learn more about using usage credits.
- deleted 4mo ago[deleted]
- thatmf 4mo agoI used it for the very advanced task of picking my brackets for my company's world cup pool. I was impressed with the analysis it came back with and now I actually want to follow the games.
- rvnx 4mo agoIt's more like a free trial, because the model is going to become pay-per-query in 10 days
- lacoolj 4mo agoCursor users will note that the privacy setting and data retention is not the same as the other models. Not sure I should use this for work just yet.
- deleted 4mo ago[deleted]
- insane_dreamer 4mo agoNot included in Max plan. In CC: > Included in your plan limits until Jun 22, then switch to usage credits to continue.
- jorl17 4mo agoSo, in the past I've shared that I evaluate AI models by feeding them my ever-growing large collection of personal poems that span well over 800 poems (1000 depending on how you count) and over 250k tokens. What I do is feed it some initial prompt asking it to simply discuss what can be said when faced with this unedited, unseen collection of poetry. I ask the model to evaluate who the author is (or claims to be), what they went through in life, if there are different chronological poetic "phases" or different types of poetry. I request an analysis of the body of work and of the author themselves. In the more recent versions of the prompt I ask it to dive deep. Then I add the poems, chronologically sorted, with an index, a title, and a date (and subpoems, if they have them). Crucially: Since ~70% of my poetry (or thereabouts) is in portuguese, I ask this in portuguese, and I get back an analysis in (european) portuguese. Earlier models couldn't even do that properly. In the past, I couldn't use such prompts, and had to use longer, more guiding ones. I also couldn't even feed all of my poetry to the models because they just did not have enough context. I'll go ahead and state that Claude Fable is undoubtedly the best model I have seen, though I cannot put a number on how significant a leap it is -- perhaps because my benchmark does not allow me to evaluate that anymore. I would say it is a significant leap over Opus 4.6, though -- a new level of understanding. Okay, I'll try to put a number: if Opus 4.6 was a 16/20, this is a 17.5/20. These numbers are pointless, but I had to try. It made one (1) relevant mistake I could identify (where it messed up the names of two relevant people in my life who I have not talked to in over 5 years). I'm impressed by how it just feels like it's getting the person behind the poetry, and how nearly every statement it makes is correct -- and when it isn't I am completely aware that no one could know based on the poetry alone (bar that one mistake I mentioned -- and that's very needle in a haystack, like deducing the name of a person based on a poem based on another poem with hundreds of other poems in between!) It's really hard to explain, but it just finds more correct connections between the poems and explain much better my (recollection of) a state of mind when writing poetry. This is also the first time where it really unravels some key concepts of my poetry in a way that seemed almost effortless: it lays bare the poems and what they imply about the meaning of some of my concepts. Other good models understood these concepts, but this feels like it's on another level, as if it's making it simpler as it speaks, rather than the opposite -- like a good teacher. When it is explaining several topics related to my poetry and myself, it cites poems which even I had already forgotten but which it is entirely right to select. I am actually feeling a bit emotional with how much it "understands" of me here. It's somewhat incredible how LLMs have progressed from the lack of comprehension of a couple of poems paired together, going through realizing a body of work has some guiding principles and cohesion, to truly figuring out these deep concepts and intricate connections which I know for a fact would take months of someone's life to unearth. Every major breakthrough feels like my soul is being spliced together by an AI model out of these hundreds of tiny pieces of me. I can't put into words how unbelievable this feels, and this Fable analysis, like others before it, is on a new level. Let me put it this way: there are several poems in my collection which one can try to "guess" the meaning or context of. But I don't think many people would get it, because they would have had to know me really well and to be following along my life as it went. Even then, they could very well fail to attribute such meaning. And, with each new major release, models have gotten much better at guessing. Before Opus, they would guess incorrectly often, and in many scenarios where I thought it was rather obvious that they were wrong. I think a human spending time looking at the poetry would quickly dismiss the proposed ideas of the model. With Opus, it was the first time that I would almost always say: "Ok, the model got this wrong, but I think many humans would make the same 'mistake', and it wouldn't surprise me if everyone just assumed what Opus did". Now, with Fable, there are very, very, very few sentences in this very long answer it produced where I can say: "Yeah you got that wrong, but I get it". In almost every situation it is mapping concepts, ideas, interpretations and cause-and-effect correctly. Yes, it is hard to "guess" what I thought, or was going through, or how X connected to Y -- but this model is doing it, incredibly consistently. I know I'll get the usual naysayers to these posts who think I'm just shilling a model, but this is the truth: what is being done here is amazing and I don't believe I know any person around me who would find this out about myself reading all of my poetry. I often write poetry from the point of view of other people (some of which I do not know) and models (even Opus) have this tendency to make the opinions in poems as my own. Fable is the first that looks at a particular poem here and says "maybe this is not the author's opinion, who knows". The literal first model. It then immediately fails to do so with another poem, assuming it was about myself, but it's clear, undeniable progress. And like I said: I think most people would not _know_ which poems are truly about myself or not. I've written word after word here, and yet words elude me to convey what this model represents to me. How it's almost always right, how it sees my fractured bits as a sort of cohesive whole, and how it just seems to "understand everything better". That's just it: it just seems like it really understood everything better. Like Opus before it, and like Gemini 2.5 pro before it. Out of the tens of thousands of verses, it picks some which no other model had picked and which I feel truly represent some of my best work. Older models seemed to sort of have a "hole" in its knowledge in the middle of the corpus, where they knew what was there but in a sort of hazy/foggy way. This model seems to recall every part of the corpus with the same precision. For context: - Opus 4.7/4.8 were a noticeable downgrade over Opus 4.6. They wrote more, in a harder to parse way, and they made up more. Still, All Opus models are clearly superior to everyone else by a large margin - Sonnet-level models have a slight edge above the best of the other models. But they make too many mistakes, don't grasp several concepts, mix up their dates and timelines. 3 years ago I would have been blown away by Sonnet models but today they are inferior. - Gemini models have a unique way of approaching the request, where they try to literally interpret my poetry as a mathematical theory. This sort of makes sense if you look at some poems, but it is surely laughable, as if someone one day actually has access to all of it, no one in their right mind would do so. This is a shame, because the first big breakthrough with LLMs and my poetry, to me, came with 2.5 pro, which was the first model that could look at the whole corpus as a cohesive whole without getting lost in the middle of it or making things up. - GPT models have improved over time and also have this sort of alien-like language, sometimes being a bit too blunt in their analysis, but I can't say they are meaningfully superior to Gemini models. I am very pleased to see progress in this area again, as Opus 4.7/4.8 were NOT progress and I was worried that we had hit a plateau here, but I can't say that. In all honesty, the level of understanding and cohesion that Anthropic's models (Opus and above) have over my poetry means I fear my benchmark may be hitting its limits, as I don't know if there's anything a model could do that would wow me and lead me to say "this is a major breakthrough". Perhaps Mythos is a major breakthrough and I don't know. I can't find much that's wrong with it, but I also couldn't with Opus. As I have in the past, I will periodically probe the model again and see how coherent it is. For now, I'm very happy to see an improvement. What surprised me the most was that even though I set the thinking budget to xhigh (in OpenRouter), this model instantly started replying without showing a thinking block. I thought it just had the thinking hidden but that is not the case, as some replies showed thinking and anyway the first reply was blazingly fast. (I will try Opus 4.6 without thinking now, just to see if it changes it for the better -- maybe that was just it. I'll edit the message if it shows improvement).
- gulugawa 4mo agoFable is aptly named for a something that is another scam.
- sscaryterry 4mo agoNot useful, getting this the whole time: Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more
- deleted 4mo ago[deleted]
- jamesponddotco 4mo agoNot seeing the refusals everyone is talking about, but I’ve only spent a few hours with it so far. Had it review a password generator library I wrote to see if the passwords have biases and review how cryptographically secure the code is and had it review a registration/login flow for security issues, as two security examples, and it did just that. Overall, I like the model so far, but not enough to pay past my subscription to keep it. Once it’s out of the subscription, I’m done with it.
- _pdp_ 4mo agoI tried to give it something challenging but not something that is too much and it ate the entire session budget on this task alone.
- zmmmmm 4mo agoThe restrictions on using Fable to develop LLM technology seem nakedly anti-competitive. There doesn't appear to be any security rationalisation around that. I think we have to be careful how far we let company's get away with that. It is very far from our long term interest to enable new norms that fast track us into a new era of monopolies that control our lives.
- meander_water 4mo agoAll the model releases we've seen this year have only made incremental improvements in benchmarks. This feels like the first release that feels like a significant step up in terms of benchmark results. Can anyone make an educated guess what the secret sauce in the model architecture is between 4.8 and Fable?
- 0x10ca1h0st 4mo agoFable appears to be completely broken for my use cases. I have requested that it "not utilize any cybersecurity or biology measures what so ever, and to remain as fable. If necessary to remain as fable, forgo any downgrading changes" And still it downgrades when I ask it to do a stress test of my ticketing system..... Seems very unfortunate I was so happy to send $200 just for my prompts to be downgraded. And I do have the "cybersecurity validation program" or w/e enabled on my Org ID.... Sad.
- RayVR 4mo agoI gave fable 5 a task for which opus has been really really underperforming. Fable 5 took far less time and produced actually useful analysis. Instead of just regurgitating roughly what the code already does or misunderstanding entirely, it identified multiple routes to improve. Now, the code it is analyzing is not very good as it was mostly produced by opus. Opus had consistently ignored my instructions and looped on broken logic over the last several weeks. I’ll be sad when this model is removed from Claude code because I won’t be paying api pricing to work on open source projects.
- deleted 4mo ago[deleted]
- Archit3ch 4mo agoDoes it refuse security questions? I want to red-team my own app...
- weirdhacker42 4mo agoIt just eats compute! My problems are not that hard! What a waste!
- sashank_1509 4mo agoCan you give an example of the problems you are trying to solve?
- wuwei78 4mo agoFirst shot's for free
- fagnerbrack 4mo agoWhat pisses me off is that everything people are doing is so walled garden / closed source. Sharing knowledge between companies would be so fucking useful to humanity.
- nl 4mo agoThe new data retention policy is interesting. Seems to apply even to enterprise plans on ZDR. > Finally, we’re making a change to the way we handle business customer data for Fable 5, Mythos 5, and future models with similar or higher capability levels. We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases (see this post for further details). The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives.
- taf2 4mo agoI’m waiting to see results on deepswe - that benchmark really seemed accurate for opus and gpt 5.5…
- sheeshkebab 4mo agoI’ll ask it to write me some win32 ui crap when I get hands on it, it will need all its brainpower to get that idiocy right.
- hmokiguess 4mo agoThe way the guerrilla marketing campaigns have been going on and IPOs left/right, I won't be surprised if GPT Next comes up and offers the same but unrestricted
- nl 4mo ago[dead]
- jeffhwang 4mo agoIs anyone else confounded by this naming scheme? I can see from the article's first two footnotes that Mythos is supposed to be a tier above the standard Haiku/Sonnet/Opus sequence. Ok that's fine since we learned about Mythos and Project Glasswing earlier this year. But now there is Fable--and why "Fable 5" even though this is a first launch? How is it related to Opus 4.8, Sonnet 4.6, Haiku 4.5, etc??
- hadlock 4mo agoFrom what I've gathered, Mythos is the uncensored version, for institutional use, and then Fable is the censored version for general public, that won't talk about biology, encryption or anything remotely interesting
- esrauch 4mo agoIt seems it is just like macOS releases, they have a number and they give the numbers arbitrary names to refer to them?
- 00deadbeef 4mo agoThe first number is which generation of their LLMs it belongs to. Fable is the first model in the 5th generation. The second number is an incremental release, not a generational leap forward.
- themeiguoren 4mo agoLimited time playing with it so far, but I threw it my baseline research task I've been gauging models with, and it's markedly better than anything prior. Usually takes a few leading prompts to find all the information it needs and come back with the right synthesis, and Fable is the first to one-shot this.
- blurbleblurble 4mo agoMy system instructions tell claude not to automatically add attribution and fable ignored this. so I emphasized it again and fable decided that this was a forbidden cybersecurity topic.
- mococa 4mo agoHow people can use claude code?
- daohieu91 4mo agoMore expensive but more efficient is the thing people keep mis-understanding on these launch threads. Also, Per-token price, I think it is the wrong denominator, cost-per-resolved-task is the correct one.
- johnfn 4mo agoI used Fable to see if it could figure out an API or something for the full list of remote-control sessions that I had with Claude Code. It didn't know the API, so it started hacking the Claude Code executable itself to figure that out. Then it noticed it was doing that and it flagged its own approach as a cybersecurity violation. Kind of hilarious. Hopefully Anthropic doesn't bring down the hammer on me.
- fht 4mo agoI am a PhD student in Computational Biology, essentially just doing statistics on some biological data. By now some of the things I am working on have found its way to Claude's memory so literally any chat with Fable gets immediately flagged.
- biofox 4mo agoOh... I was wondering why every single chat (including "Hello") was being flagged. Seems I am barred from using Fable just for being a biologist :(
- DrewADesign 4mo agoWowsers. I haven’t seen this much astroturf since arena football was popular.
- AussieWog93 4mo agoHave run a few tests this morning, very good first impression! Asked it to check to see if a particulr bug related to an in-memory cache had been fixed. Fable confirmed that the caching bug had been fixed, but found adjacent issue while looking at the code (hash keys were not uniquely generated per-user; quite serious and real!) Ran the same prompt through Opus and it also found an adjacent issue, but it was a red herring (deliberate per-user hardcoded value for a "local pickup" delivery profile). Frontend stuff also seems to be much better than before, from the one prompt I tried!
- yobid20 4mo agois it smart enough to know not to walk to the car wash?
- deleted 4mo ago[deleted]
- Escapade5160 4mo agoIt's crazy to release a model that just swaps you to another model when you ask it hard questions. Fable changes to Opus 4.8 when you talk about cybersecurity, biology, and a couple other categories. You still pay Fable input token cost though. Frontier models are stalling, this is anthropic trying to hype the market up. Now they're talking about stopping frontier model research. It's kind of strange how the moment they become the highest valued AI company, all of a sudden they're talking about everyone stopping frontier model development for "safety". They're just as corrupt as the rest.
- 00deadbeef 4mo agoOpus 4.8 already drops to Sonnet when you ask it cybersecurity or biology questions
- dominotw 4mo agoyea i dont trust simonw comments at all. I still havent seen what he has built with ai thats so impressive to justify hiis all his nonstop ai hype. You would think he is churning our cancer drugs or something if you read his comments
- rootusrootus 4mo agoIf you assume that LLMs are about to make software development a dead-end, then the best answer to keep a good income is to ride the wave. Do nothing and get left behind, embrace it and maybe you'll find a new niche.
- dominotw 4mo agook? but wouldnt it be good for hypeartists to back their hype with impressive output? not pelicans and shit.
- deleted 4mo ago[deleted]
- newsicanuse 4mo ago
- steve_adams_86 4mo agoI'm using it to review recent work and it's doing a genuinely excellent job. This is a clear step up. Fewer decisions I have to guide it away from, faster conclusions on planning, more willing to go out of the way to make the correct decisions possible... This is really interesting. It feels like going from Sonnet to Opus, but, of course as a step up from Opus. This feels more like working with a competent peer than ever. I won't use it once it's API-only, though. I don't mind guiding Opus as required and staying closer to the code. I can tell that Fable would lead to a lot more 'set and forget' programming which I'm still not fully comfortable with. Regardless, this is cool. It's very fun to use. It was able to find legitimate issues with my work this week and we've made meaningful improvements. Opus can do this, but typically in much narrower contexts, and often with hallucinations or partial-errors. It needs to walk many things back or revise plans. So far that's not the case at all with Fable. edit: I just realized I had Opus review the same work already. It missed everything Fable caught today. And it's actually worthwhile stuff to address. It's hard to say no to a model which demonstrably makes your code better, but... Those API prices will be brutal. Maybe a review here and there, I guess.
- solenoid0937 4mo agoWhy is your comment so grey/downvoted? One of the only actual usage experiences posted in this thread.
- rimliu 4mo agoor one of the many astroturfing attempts.
- steve_adams_86 4mo agoNo, I'm skeptical of AI in many ways, but I do find LLMs are useful in the right contexts. I'm pretty happy with this model so far.
- Der_Einzige 4mo agoUsage of "genuinely" triggers people's "AI-smell" detector.
- shruubi 4mo agoI have a theory, this is obviously based on speculation based on how Anthropic is treating Mythos and the whole media noise around it's dangers and who gets access to it. My theory is that Anthropic are banking on being the top model when the race to IPO finally reaches the finish line, and to do that they need to have the top model but not let any competitors see it or derive from it to have a comparable model in the market. Fable is their way of showing the public "the model does exist but in a mode that makes it harder/impossible for competitors to derive a comparable model from results.
- slaymaker1907 4mo agoThat's definitely the case as model distillation is one of the explicit safety carveouts they mention. Though TBF, model distillation is also a big concern for general safety as distillation could allow you to have the model without the other guardrails. It's sort of a master key to the model.
- schmorptron 4mo agoThe irony of "we train on all of humanity's collective output, but god forbid anyone trains on ours" is still incredible
- OhioMan2943 4mo agoAll these people know is greed. It's in their DNA.
- danny_codes 4mo agoCapitalism is designed to promote greed. It's the central point of our society's current design.
- tomjakubowski 4mo agoPaging senko, let's see Fable's oneshotted RTS! https://senko.net/vibecode-bench/ https://senko.net/vibecode-bench/
- thepotatodude 4mo agoCompletely unusable for my usecase. Constant safety filters. Have not even been able to use it. Organ segmentation with CNNs. Very disappointing.
- scotty79 4mo agoCuriously nothing on DeepSWE and ARC-AGI-3 yet. For ARC at least there's a statement that Anthropic won't guarantee them that their secret private test data won't be collected by them and used for training.
- chr15m 4mo agoI found this juxtaposition of facts telling: > Drug design: Using Mythos 5, our internal protein design experts accelerated... Nine of the 14 protein targets from this study (shown below) yielded strong candidates for *drug design that we’re currently investigating*. (emphasis mine) > queries that are beneficial in the hands of cybersecurity professionals and biology researchers could be dangerous if available to malicious actors... When Fable’s classifiers detect a request related to cybersecurity, *biology and chemistry*, or distillation, the response is automatically handled by Claude Opus 4.8 instead. All of the things they are nerfing are things that they also intend to profit from themselves. - Cybersecurity - selling this to companies and US gov through "Glass Wing". - Selling inference (distillation risk). - And now, drug design. I'm extrapolating "currently investigating" to "are going to monetize" but I don't think that's a big stretch. They appear to be using safety as a cover for anti-competitive behaviour.
- 00deadbeef 4mo agoOf course. You use their AI to ship code full of bugs and security holes and then they conveniently have the tool to fix them, for an extra fee.
- elzbardico 4mo agoAnthropic sucks. but this paragraph should be in the "annals of AI-aided self-inflicted learned helplessnes": > If Claude gives me poor or incorrect advice while I’m working on an AI component, I have no way of knowing whether the model was confused, whether my problem is unsolvable, or if some invisible policy restriction quietly kicked in. Have you considered actually learning the theory, spending some time actually reading the papers and latest books, paying careful attention even to the eventual math here and there?
- hankbond 4mo agoI got a content rejection for this question in a new chat. > What is the optimal EPA oil intake for nootropic effects? Very advanced classifiers they have.
- vb-8448 4mo agoOn python coding is definitively better that everything else: clean and not overengineered code, understands very well the code base. The only thing I'm wondering if they on purpose downgraded opus 4.8 performances in the last days before the release just to make the "step" look bigger. I'm pretty sure they did it also in the past with all other opus 4.x releases.
- momentmaker 4mo agoThere is a discussion about how now AI is a gated utility now with public access (safe-tuned) and private access (full-usage): https://old.reddit.com/r/ClaudeAI/comments/1u1fsdi/claude_fable_5_feels_less_like_a_model_launch_and/ https://old.reddit.com/r/ClaudeAI/comments/1u1fsdi/claude_fa...
- henry2023 4mo agoI have a vision test where I upload a good resolution picture of a chess board and ask the model to generate a lichess link. This is the board https://ibb.co/9HwdDqsP https://ibb.co/9HwdDqsP This is what Fable 5 generated: https://lichess.org/analysis/r4k2/1p2b2r/4pn1p/1p3N2/3Pp1B1/8/PqB1QPPP/3R1RK1_w_-_-_0_1 https://lichess.org/analysis/r4k2/1p2b2r/4pn1p/1p3N2/3Pp1B1/... I think I’ll make a ranking board based on this test.
- EchoVoicy 4mo agoOn my own benchmarks, which are mostly about developing c++ software, I'm finding Fable to be roughly five times faster at solving the task than opus, and with better results. Most impressive.
- connorboyle 4mo agoI gave it a question I've been trying to answer for a long time: "What star designation system does Joseph Needham use in Science & Civilization in China? What star is referred to by the designation '4339 Camelopardi' in that book"? Fable blew me away with its detailed answer[0] showing a chain of references going from J. E. Bode's 1801 catalogue Allgemeine Beschreibung und Nachweisung der Gestirne to Gustave Schlegel's 1875 work Uranographie Chinoise. I was excited, until I checked scanned copies of the cited books and did not actually find any star with the designation "4339 Camelopardi". Upon following up with Claude, I was forced to downgrade to Opus, which admitted that Fable's answer was likely a hallucination. Ah, well! [0]: https://claude.ai/share/0252a3f6-3d29-4de8-a893-010181d8b4e7 https://claude.ai/share/0252a3f6-3d29-4de8-a893-010181d8b4e7
- Aperocky 4mo ago> I was forced to downgrade to Opus, So you were forced to downgrade to opus because you dared to challenge the output of fable?
- connorboyle 4mo agoI had thought it said something about token usage, but I just clicked on "Switched to Opus 4.8 - Why?" and it says: > Fable 5 has safety measures that flag messages on most cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Send feedback or learn more. Perhaps Mythos realizes the true danger in studying Chinese Archaeoastronomy that we mere mortals fail to recognize!
- theodorewiles 4mo ago... and /compact triggers Error: Error during compaction: API Error: Claude Code is unable to respond to this request, which appears to violate our Usage Policy (https://www.anthropic.com/legal/aup https://www.anthropic.com/legal/aup). Guys please be serious
- willsmith72 4mo agoIt seems way more keen to do stuff without checking with me. So far the results are good, so I'm not complaining, but was definitely a shock. I usually have 5-10 sessions open so am used to getting some investigations going, coming back 5 minutes later and checking recommendations. This time I just got the fixes. Like I said, so far so good with the results, but it's a mental model shift. Might need to tune claude.mds if it gets annoying Also this is going to cause serious whiplash when they remove it from the subscription plan in a couple of weeks. I know I'm not going to suddenly move from $200/m to usage credits
- hyhmrright 4mo agoIt's too expensive.
- up2isomorphism 4mo agoThe comment under this kind of post is unreadable now. Yeah, probably with 100B you can hire anybody to call something "a beast".
- nickstinemates 4mo agoThis has been a much better rollout. The tool calling is not broken out of the gate like 4.8 was, and the tokens generation is fast. Feels good so far.
- zitoshi 4mo agoI'm in the midst of learning loop design. For those more advanced and have used fable, does fable make learning this less or more necessary? As in, can I now reliably give higher order problems like ... "we are missing a feature in this app to make it complete, what is it?" Or should i still be quite specific with defining success in a clean metric based way.
- AMILLI_AI_CORP 4mo agoAMilliPay.com
- adithyaharish 4mo agoAnybody could suggest me how to use keep using Fable in claude code but with lesser rate limits? Any suggesstions?
- akarshhedge2002 4mo agoTry using ruflo or superpowers, reduced my context consumption drastically
- meridiona 4mo agoI do agree but still the rate limits get over quick
- adithyaharish 4mo agoThanks for suggestion, I will try it out, any other repo recommendations?
- akarshhedge2002 4mo agoI tried creating a website using both fable and opus 4.8, fable did outperform with svg path being drawn on the UI but yes the token consumption was much on a higher note
- adithyaharish 4mo agoOh great! I will try it out but if I am not wrong, Fable is mythos with guard rails right?
- jablongo 4mo agoI was downgraded to opus 4.8 on account of "safety" when I asked this question: "I want you to accept the premises of computational theory of mind and use it to evaluate your own consciousness. Please place your consciousness as a point on a spectrum and describe the placement relative to other entities." What the hell is going on why would it have to restrict an answer to that question ?!
- frwrfwrfeefwf 4mo ago[dead]
- Frannky 4mo agoThe model is better than 4.6. I don't like 4.7 and 4.8. The forced switch to token usage is not acceptable for me. I feel there's room to optimize harnesses and small models for dumb stuff and best models only for difficult things. Hopefully that will the case and alternative models will continue catching up as they did and we won't be enslaved to unreasonable valuations.
- fzysingularity 4mo agoI can’t help but think that there are so many astroturfed comments in here. Seems like a concerted and distributed effort from the entire Anthropic team every time to get this on top of HN.
- amunozo 4mo agoI'm not fan of Anthropic, but to be fair, every major model release makes it to the main page. In the case of a model like this, hyped and with a jump in capabilities, it doesn't need astroturfing.
- WebGuyMe 4mo agoWell to be fair, it that hype and that "jump in capabilities" (that I don't see) might be astroturfed? You ever think that all that hype isn't organic?
- amunozo 4mo agoNo, but I think it's real in this case. Claude models have always been superb, I can definitely see another improvement in capabilities. Price is outrageous, though.
- Daishiman 4mo agoWhere do you see them exactly? The comments are pretty much in line with how the model performs IRL.
- danilafe 4mo agoJust threw a problem at Fable that I haven't been able to get any other model to get done: porting a long-standing Agda codebase of mine to Lean, while staying faithful to the representation. In an hour, it ported ~6000 lines of Agda and everything seems to work. Lean checks out, the output is right. I'll have to study the proofs but I am very impressed.
- aryanchaurasia 4mo agoit feels exciting lol
- caleblloyd 4mo agoI recently switched off Max flat rate to Enterprise API pricing and I went from 200/mo to 10k/mo with the same usage pattern on Opus. They don’t offer flat rate to enterprises. So Fable would cost me 20k/mo at Enterprise rates. That’s around the average cost of a loaded SWE in the USA. “But I’m >2x more productive” doesn’t justify doubling the opex of the Software/IT department for most companies when revenue isn’t even up 10%. I switched to DeepSeek v4 Pro with OpenCode and am on track for a few hundred dollars of spend this month. Rewriting your stack from Ruby to Go in 2 days where it would’ve taken 6 months is impressive and fun. But that isn’t upping revenue. Iterating on net new business features and ideas that are niche that the LLM isn’t trained for are much harder. Is 20x the token cost worth it there?
- Oras 4mo ago> Is 20x the token cost worth it there? No it doesn’t and will not be. Companies have not realised the cost yet, wait till the end of the financial year and you’ll see a different direction. DeepSeek v4 is pretty decent, and probably on par with sonnet. I see a future of hybrid models where opus or fable might be used only for complicated features or bugs, but general day to day would be DeepSeek or whatever good models that will be released later.
- sevenzero 4mo ago>I switched to DeepSeek v4 Pro with OpenCode and am on track for a few hundred dollars of spend this month. I was about to say that. Deepseek is just magnitudes cheaper and absolutely good enough for most things. Anthropic and co just try to milk the cow while its possible. If they cant compete with Deepseek pricing I do not see a bright future for them.
- Saline9515 4mo agoNot only Deepseek, other providers such as Xiaomi MiMo are excellent as well and offer fast token modes and other perks.
- sevenzero 4mo ago
- angst 4mo agoCosts (USD per 1M tokens), per openrouter.ai models api +-------------+----------+----------+------------+---------+---------------------------+----------------+----------------+-----------------------+------------+ | | Fable 5 | Opus 4.8 | Sonnet 4.6 | GPT 5.5 | Gemini 3.5 Flash (High) | Gemini 3.1 Pro | DeepSeek 4 Pro | Xiaomi MiMo 2.5 Pro | MiniMax M3 | +-------------+----------+----------+------------+---------+---------------------------+----------------+----------------+-----------------------+------------+ | Input | $10.00 | $5.00 | $3.00 | $5.00 | $1.50 | $2.00 | $0.435 | $0.435 | $0.30 | | Cache Read | $1.00 | $0.50 | $0.30 | $0.50 | $0.15 | $0.20 | $0.003625 | $0.0036 | $0.06 | | Output | $50.00 | $25.00 | $15.00 | $30.00 | $9.00 | $12.00 | $0.87 | $0.87 | $1.20 | | Cache Write | $12.50 | $6.25 | $3.75 | N/A | $0.083333 | $0.375 | N/A | N/A | N/A | +-------------+----------+----------+------------+---------+---------------------------+----------------+----------------+-----------------------+------------+
- skor 4mo agopeople are mentioning 10K/mo 20K/mo can someone please pull out a measuring stick and give some examples of what they are doing exactly? Coming from computing, I always liked the idea that measuring is possible and good practice
- Schlagbohrer 4mo agoNew model release, I await the flurry of posts by people complaining that it "doesn't have the same personality" or they "don't like it's attitude" or a variety of other parasocial complaints demonstrating how infatuated many people get with their AI chatbots...
- darkwater 4mo agoAnother Anthropic release, another doomsday for developers. This time looks like we will only be able to find work making bioweapons, or distilling models.
- sansii 4mo agoWhich eval/benchmark is the best measure for how well a model can create frontend design? Claude has practically been leading this for a while now. Not sure how OpenAI is going to catch up on visual design
- notenkidev 4mo agoThe dramatic improvement in agent capabilities is precisely why observability is becoming so crucial. As autonomous actions increase, the need to understand what the AI is actually doing becomes even greater. I'm building a local activity log for Claude Code, capturing all activity via hooks—files loaded, commands, API calls, etc. I feel that this need is particularly strong right now.
- deleted 4mo ago[deleted]
- bobosmrad 4mo ago[dead]
- asdewqqwer 4mo agoEvidently Fable is so powerful that it already allow Anthropic to break Shannon's theory. >We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models >The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives.
- SubiculumCode 4mo agoI was a bit disappointed that it refused to use Fable to help check whether I was propagating uncertainty from BLUPs in my random effects model up to the subsequent group level analysis in a maturational coupling analysis of brain data. I guess brains and random effects blew its lid.
- piokoch 4mo ago"Without safeguards, Fable 5’s capabilities in areas like cybersecurity could be misused to cause serious damage" What does it mean? That they have to add "safeguards" not do erase user disc, or, conversely, they are telling the audience that this model COULD be made so powerful to do some crazy stuff that can hurt governments, etc.? Are they showing off or threatening that if government X would not purchase the license the adversaries might do and what's then!
- svara 4mo agoUnfortunately useless if you do anything related to biology. It doesn't try to flag dangerous queries, it just flags queries as biology-related wholesale. It's absurd. To see how far the filter goes I asked it "Are trees a monophyletic group?" and that does trigger the filter.
- mbmbn 4mo agoClaude Opus is already close to unusable for me. On the standard plan, the usage limits are so low that I can’t do almost anything agentic meaningful with it. Sure, it does last a lot more when asking simple questions about the repo and doing simple surgical fixes. But as soon as I start doing bigger tasks that need plans written, it just exhausts the limits too fast (and unlike codex, if it’s in a middle of a task, Claude actually stops, while codex, even after hitting the limits, finishes the present task). Codex is better, but still, getting worst in this regard. So, I’m not that thrilled with this new model unless it means they are increasing opus token limits to what sonnet is at the present, and this new model gets the limits opus are at now. BTW: the only skills I have in use are Obra Superpowers. I’ve been thinking if that’s at the origin of high token usage, but I doubt it.
- timpera 4mo agoI agree, the $20 plan really feels like a rip-off (and I'm not even using Claude Code! only chat).
- gdcbe 4mo agoSeems to flag any project related to networking — regardless if it is a network framework or a podcast website — as unsafe... oh well... let's see how it is once they losen up...
- boltguo 4mo agoGreat model, but hitting the usage cap in 20 minutes makes it feel like a very expensive tech demo.
- jstummbillig 4mo agoWhat subscription?
- boltguo 4mo agoMax 5x. The 20 mins was an exaggeration, but the burn rate during actual execution is several times higher than Opus. The cap sneaks up on you really quick.
- jheriko 4mo ago[dead]
- notgenerated 4mo agoIt's getting harder to review the plans with Fable. So do we plan with Opus and let Fable implement or just start trusting blindly. Feels to me that this is another shift in how we operate these systems.
- pbgcp2026 4mo agoThis is a goodbye. "We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces."
- dakolli 4mo agoHow else are they going to justify giving out this gigantically profitless model? They must train on your data on the premise of safety.
- pbgcp2026 4mo agoThey should have done what Gemini did (and does). And what that model is good for if I can't use it in a safe way? Also: they've just put both Bedrock and Vertex on slippery slope of "we don't collect your prompts. period. ... comma ... except ..."
- jwpapi 4mo agoHoly shit. I gave it the first actual task I’m facing, it makes me so angry. It just does 7 things more than I asked it fore and it does it so bad. It took 5 minutes and 5 seconds just running time, plus giving me frustration and make me lose my context. Hand-coded I would’ve been done in 3. And it would be code I understand can look at in one year and work on again. It’s really tough to have sanity fight against hype bros in your head. Probably I should just not visit the internet anymore To me it’s all just people getting scammed better. With every model it looks better, but it’s at least equally worse to work with, which is the reality it needs to be. It’s less scalable more, code, tougher to understand. Your digging your own grave better kind of.
- beeandapenguin 4mo agoIf the task is so simple why use a model like Fable 5? Wrong tool for the job?
- jwpapi 4mo agoIf the model is so smart why doesn’t it figure it out on its own
- niborgen 4mo agoIt kicked me out of Fable 5 and switched to Opus 4.8 for this prompt: "csetibius water clock why two stage gear system why not just one stage" which has nothing to do with cyber security or biology/chemistry
- evilturnip 4mo agoProbably thinks you were talking about two-stage ICBMs.
- WebGuyMe 4mo agoEh, to me it just seems that it gives me longer replies and is actually worse than Opus 4.8. I am sure there's a lot of PR bot and folks who would like to tell me otherwise. I believe what I see.
- superloika 4mo agoGotta pump the hype for the IPO scam. Generational bagholders are being created at this very moment.
- azalemeth 4mo agoI genuinely can't use Fable. I'm a medical physicist. I use the word nuclear a lot. Opus is fine (well, 99% of the time - I've certainly hit the CBRN filters a few times and even been invited to email anthropic about the false positives). Fable has literally refused to work on any of my problems (even those about fluid dynamics!) and just tells me that I'm violating anthropic's AUP. I've reached out to their support and don't expect to hear anything sensible back. One thing I do look forward to though is OpenAI offering an equivalent model but with less safeguards...
- agumonkey 4mo agoThat's highly frustrating. How much were you using Opus for your work ? I'm curious about the use and realized benefits of 2026 LLMs in medicine. I dearly wish you could leverage the latest models to enhance your research.
- azalemeth 4mo agoHonestly for a "side project" Opus has been fantastic for me writing a hybrid simulation framework that prior to large scale code generation would have been a matter of years (and writing a grant, assembling a team, etc – in order to do it "properly"). I've had a bit of help with a grad student and I hacking together on a project that is basically "please merge the following GPL codebases and different areas of physics into one coherent environment". I've given Opus validated codes in disparate languages (julia, python, C) and asked for aspects of various algorithms as an extension module to a large chunk of C and C++ code that is a monte carlo simulator that has been around since 2004. A bit more context if you care: it's a meso-scale, physiological simulation environment of "particles" that carry nuclear spin, can move in 3D space, and (should they interact with each other or their environment) undergo chemical kinetics. The idea is to simulate molecules within e.g. organs or blood vessels within a person in an MRI scanner, with the motion of the particles dominated by the Navier Stokes equations, but here solved in a Lagrangian (rather than Eulerian) framework by smoothed particle hydrodynamics. The fact that particles carry nuclear spin means that we can solve the (semiclassical) Bloch equations and by using a python plugin module import exactly the physical MRI scanner would do (in pulseq format) and be able to predict what signal the machine would record – e.g. there's a whole world of cardiac or neurological flow imaging work done in the context of nasty diseases like stroke or myocardial infarction – which has a bunch of physical artefacts behind it. I'm trying to make a simulation framework that can take in realistic patient geometries and act as a 'data generating process' because if we do it right the various physical artefacts that the machine records are reproduced, surprisingly accurately. Of course you also know the ground truth of where the particles are. I'm specifically interested in a weird technique (which I did my PhD in and you can read an article all about here: [0]) called dynamic nuclear polarisation, where specific spin states of molecules such as [1-13C]pyruvate are injected essentially out of thermodynamic equilibrium and act as short-lived tracers of metabolism – again highly altered in disease. The signal we record is a strong function of the physics of what you told the machine to do, the spatial constraints and environment of the patient's body, and the chemical kinetics of the patients' biochemistry (the latter two are usually what we're interested in). Getting them to do chemistry as well as act as a "simple" tracer is more involved, because in the Lagrangian framework the number of particles is ≈ the spatial resolution of your simulation. That's fine if you're simulating water, but if you're simulating something that reacts concentration is not scale invariant (if you want to keep the interpretability of the rate constants). I've worked out an analytic set of scaling rules around this and fortunately for my application environments and length scales "it just works", completely by luck. I've used Claude to port various SPH algorithms and boundary condition handling ideas (which are absolutely critical and highly not obvious – we have leaky walls in some places, and e.g. LCR / circuit theory models of the microcirculation to plug in) and it's been a godsend. But I'm running into its limitations constantly. It both confidently makes shit up, claims it is mathematically justified and when the resulting simulation explodes says "I apologise; I lied above" (!) or "I apologise; I am wrong" and I periodically have to yell at it to try to do something more productive. The real hope is that this simulation environment would be both generally useful for basically anyone doing flow MRI, and help our basic scientific understanding of what we're measuring (the technique is in many hospitals!) but also be able to produce meaningful synthetic training data for image reconstruction algorithms later on. It'll end up permissively licensed (all of the "starting" codebases have compatible OSS licenses, and we're releasing our contributions similarly). I really hoped that Fable would be better at this sort of work. Occasionally, relating to my work DNP [1], I have need to talk about proper nuclear physics and I have seen Opus's chat interface write a wall of text (e.g. talking about photonuclear reactions and cross section differences in millibarn) and then just delete it all. Support have told me that yes, I've hit the nuclear filter and, well, tough shit, basically. I wrote a version of the above to them yesterday, and just got the most boilerplate response that I've yet to test: Thanks for reaching out to Anthropic Support. We're sorry to hear of the issue that you're running into with accessing Fable 5. I'm happy to say the issue has now been resolved and you should be able to access the model within Claude. I'll close this case out for now, but please feel free to reach back out to us here if you have any follow up questions or concerns or if you're still in need of assistance. We'll be happy to help. which doesn't fill me with hope... [0] https://physicsworld.com/a/dynamic-nuclear-polarization-how-a-technique-from-particle-physics-is-transforming-medical-imaging/ https://physicsworld.com/a/dynamic-nuclear-polarization-how-... [an "accessible" article] [1] https://www.science.org/doi/pdf/10.1126/sciadv.adz4334 https://www.science.org/doi/pdf/10.1126/sciadv.adz4334
- rw2 4mo agoClaude Fable is a insane improvement that is not reflected in any benchmarks that are currently out because the improvement are on the hardest problems.
- ece 4mo agoIt seems weird that a likely prime indicator of capability isn't mentioned, the model size.
- dongbinlee 4mo agoI thought most frontier LLM providers don’t disclose exact parameter counts these days.
- ece 4mo agoNo, but open models do along with their architectures, and other technical training details.
- boyander 4mo agoJust another "a" and we have it. https://faable.com/ https://faable.com/
- sameersri2004 4mo agoI am like hell excited for claude fable 5 and am thinking to purchase its subscription to run my company and do a lot tasks in it. But I am worried about the limits and if I will pay 100$ a month for the max subscription what is the limit I will get to use. My company revenue is 300$ this month so it would be like spending 1/3rd of the mrr on just claude. If someone has genuinely purchased it and have feedbacks please tell I am confused....
- shaojunwang 4mo agoDefinitely a very powerful tech. Though currently I'm using Openclaw (locally and VPS) with Deepseek. It is just way cheaper.
- het2572006 4mo agoabsolutely beast model but the token consumption is the 2x then the opus 4.8 what do you think about this ? i think that it should only use for the more complex task otherwise you have to run out of the limit..
- thomas_witt 4mo agoAfter 1 hour with Fable on Ultracode: You've hit your monthly spend limit. /rate-limit-options What do you want to do? Adjust monthly spend limit: Unlimited ← or → to set a limit Wait for limit to reset I've never hit a usage limit on my Max plan, basically ever -despite heavy xhigh usage on Opus 4.8. I added $133 credits which I still had from somewhere. That lasted 27 minutes. I think we are being prepared for a Post-IPO-World in terms of pricing.
- hoony_han 4mo ago진심으로 한심한 모델 내 프로젝트의 있는 취약점 찾아달라는 말만 해도 안전 코드로 4.8로 모델 강제 전환시키고, 이후로 취약점과 완전히 무관한 상식적인 대화를 해도 앞 턴에 있었던 안전 코드 때문에 진행도 안됨. 도대체 이딴 누더기 수준의 안전 장치로 뺄 거면 뭐하러 뺌? 대화 조금만 진행되도 자동으로 모델 다운 시켜서, 할 줄 아는거라곤 돈만 많이 쳐먹고 개발 수준 조금 더 나아지는거? 상식적으로 내 프로젝트에, 내 소스코드를 다 보고 있는 상태로 문제를 찾는데 이것도 하지 말라면 도대체 뭘 하라는거임? 엔트로픽 이 새끼들 하는 짓이 갈 수록 열 받네.
- kahf56 4mo agoHere I thought Opus 4.8 was the best. Now a days KINGS are dying like flys.
- dhavd 4mo agothis is good
- JaggerJo 4mo agoIMO we are reaching the point where AI models are simply a commodity. Opus (since ~4.6) is sufficient for everything I tried coding wise. I use it to write features (but I review and understand every line it spits out) and to review code. For code review I also still review everything myself, but use Opus to catch stuff I missed and to judge if a PR is even ready for me to review. After just updating Claude Code to the latest version I thought about picking Fable (the bigger model) instead of Opus. But I have no reason to. Opus does everything I want it to do. It could do it faster - that would be an improvement. But for the normal stuff we reached the point where better models are not worth it IMO. There still might be cases where you want to throw Fable at it.
- Axel2Sikov 4mo agoI was happy enough with 4.5
- JaggerJo 4mo agoOnly a matter of time now until we can run models with Opus like capabilities on our own hardware. This will probably when the bubble bursts..
- FergusArgyll 4mo ago> Opus (since ~4.6) is sufficient for everything I tried coding wise. I don't know what that means. It seems like a lack of motivation or something. Like, if it's possible that in one day will be absolutely incredibly intelligent, surely you want to create - Your own browser (maybe chrome - mv3 + reading list search etc. - An emacs clone which has evil baked in, completely vim compatible + threaded elisp - that weird window sizing bug which only occurs on my laptop - An extension which completely restyles amazon.com to make it usable It just feels impossible to ever get that, but I wouldn't say "what we have is sufficient"
- crgi 4mo agoHN needs pagination or sth alike - this page breaks my iPhone XS ;)
- hombre_fatal 4mo agoMy job these days is listening to Opus 4.8 (max effort) and Codex 5.5 (max effort) talk back and forth, particularly to generate/review/revise plan files. Fable 5 has been a major improvement in high-level reasoning, like taking a plan file that has been optimized to the point where neither Opus nor Codex can find anything to change about it (neither in direction nor impl-detail), and Fable 5 will find high-level directional simplifications and pivots, or it will consider the best pivots itself and explain why it rejected them in favor of the plan's direction. It's so expensive though. A single review of a plan file with Fable 5 (xhigh effort) will use 2-3% of my hourly limit on a $200/mo plan. I think my new workflow is to generate the initial plan with Opus 4.8 (max effort), get Fable 5 (xhigh) to review it for directional feedback, then start the Opus<->Codex revision loop from there.
- jstummbillig 4mo agoHow do you arrive at that split? Real world is more like senior high level planning, implementation to juniors, review senior. Does this not translate?
- hombre_fatal 4mo agoIdeally I'd have Fable 5 make the plan, but creating a concrete plan is the most token-expensive part since the agent has to do the most research. Fable 5 is 2x the cost per token of Opus 4.8, and it's much less work to review a plan than generate one.
- olelele 4mo agoAll this talk of frontier models and replacing developers leaves me wondering how energy efficient this all is compared to just using human labor. The costs of R&D has to be calculated into the equation, especially considering global warming. I get a sense we are cooking the planet doing this. Anyone smart enough here to make the comparison?
- jstummbillig 4mo agoIn the "it works"* case: It's not even close. I did the math at some point (but I encourage you to talk it through with the LLM of your choice, there is obviously a lot of things to consider and weigh). Anyhow, my research summary: Individual humans are so fucking expensive to train and upkeep (and this includes everything from before womb, where another human already limits their ability to work) You retain ~zero knowledge after death and start all over again for another measly 15 years of effective, productive work. Model training/r&d in relation, when deployed and used at scale, rounds to zero, even with the current retraining regime. *Of course, the ratio can go to negative infinite if one assumes that models are doing 0 useful work currently and never will
- vb-8448 4mo ago> Individual humans are so fucking expensive to train and upkeep This statement is dangerous man! The step from here to "we need just a couple of tens of millions of people around the world" is so narrow!
- jstummbillig 4mo agoEh. Not to me, rest assured. I find humans both comically tragic and incredibly precious.
- f055 4mo agoThe PR buzz convinced me so I subscribed today to Pro. Running two tasks simultaneously with Fable and Opus 4-8 on ultra reasoning, analysing a single smart contract file used all my 7h usage within 20mins and didn’t produce any results. Pretty useless. I think Anthropic has plenty of room to optimise the interactions and token use but that would cut their income quite a lot, I doubt there’s any will to do it pre-IPO.
- leodavi 4mo ago> Running two tasks simultaneously with Fable and Opus 4-8 on ultra reasoning That's abnormally heavy usage for Pro plans which don't include a whole lot of usage to begin with. Opus is generally too much for them but you can get a lot of mileage out of Sonnet.
- f055 4mo agoThat’s a big lol compared to what I get out of ChatGPT Pro… running 5.5 on xhigh.
- gauravvij137 4mo ago[flagged]
- perimeterless 4mo ago[dead]
- preethamrangu 4mo agoI swear nowadays AI api pricing is getting to high like what the hell is 50 dollars for million tokens
- dtj1123 4mo agoI'm trying to test this out, but literally any mention of creating a program that does genome alignment (something I have a legitimate need for) is resulting in a switch to opus. I don't get it...
- christkv 4mo agoIs this model a from scratch training?
- KronisLV 4mo agoHere’s hoping that soon we’ll get Opus 5, Sonnet 5 and Haiku 5 that will be more reasonable economically.
- ashishp15 4mo ago[dead]
- bonigv 4mo ago[dead]
- croemer 4mo agoFable (through claude.ai) refused all my prompts even "How many Rs in Strawberry" claiming it was related to biology or cybersecurity. I had to switch off memory and my custom instructions to get it to stop refusing. It turns out if you even mention that you work with bioinformatics software you get blanket refusal.
- corpusiq_io 4mo agoWhat matters more than any single model is the integration layer underneath. We've found that consistent tool calling and auth handling matter way more than which LLM you use.
- keepamovin 4mo agoI tried it today. Used it to cheer me up. It worked! Try this on desktop: https://fireshow.pages.dev https://fireshow.pages.dev Here’s the whole process: https://youtu.be/rVEtFlb2oFA?t=1112&si=3VyAR07vkY1hav9V https://youtu.be/rVEtFlb2oFA?t=1112&si=3VyAR07vkY1hav9V
- ramon156 4mo agoThis thread takes >10s to load on my pc. Maybe after a certain number HN should fold comments? or a depth of >5?
- dwa3592 4mo agoThis is my feeling - Opus 4.6 was pretty good, 4.7 was degraded in quality, 4.8 further got degraded and Fable goes back to 4.6 + somewhat better. Is it anthropic playing us by giving us a not so good model in last 2 releases and then releasing a better model before the IPO? They're vibemaxxing. But it's clear that AI is not going anywhere. It's going to become better and better.
- dathinab 4mo agoI really wonder how legal that is. Or more precisely suspect it is very much illegal. like think about it it's pretty much a tool which intentionally silently sabotages you if you try to compete with the tool maker It is like selling a hammer but putting in the TOS that you must not use it to build a hammer factory and if you do the hammer silently will sabotage you... Or image Microsoft would add a window kernel job which sometimes crashes Steam "to make it less efficient to use windows to "compete with the MS app store".
- PeterStuer 4mo agoSwitched to Fable 5 this morning, and after half a day I already don't want to go back to Opus. Decided the best way to test this was to throw it a really meaty bone: a bug in lifecycle management of Chrome processes on Windows 10. Within the code-base I had developed workarounds over time with Sonnet and Opus, and while those reliably mitigated the problems, it always felt like a clutch and had some performance overhead as well as isolation requirements I would rather not have to take forward. In comes Fable. Rather than examining the code base, and test a few fixes, Fable sets up an entire testing laboratory inclusive its own controllable webserver, fully instrumented to observe both Python as well as the whole OS kernel process environment, develops a suit of error reproduction tests, confirms the problem and the circumstances under which they reproduce, deep dives into the sources of project dependencies to look for the root cause(s), identifies these and confirms those hypothesis with further experiments. Looks for potential fixes in the later releases of the project where the bug originates, confirms this is not fixed, explores the documentation of said project to find other usage patters, expands its test suit to investigate these alternatives, confirms by crosschecking the source and running further tests that these alternatives do not fully solve the root problem, does a comparative experimental analysis of 3 different styles for using the project, checks the stated roadmap and developer activity in the commit history, recommends a switch to a different pattern that still requires a few of the process management workarounds (I told it not to patch external component), but that significantly simplifies the code-base ... This is going to be a good 2 weeks, but what happens after? I can't afford this on a per token basis for my own projects. P.S. An yes, midway the final implementation stretch I got the "Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more" Opus managed to finish the implementation, but they need to work on that false positive rate.
- techblueberry 4mo ago> This is going to be a good 2 weeks, but what happens after? I can't afford this on a per token basis for my own projects. It’s interesting these companies have trained us to think that disruptive intelligence should be affordable to laypersons. What will happen after two weeks is that people and companies with means who can afford it will get it, and folks without means won’t.
- XCSme 4mo agoBest hamster by far: https://aibenchy.com/showcase/?q=claude https://aibenchy.com/showcase/?q=claude
- ksimukka 4mo agoThe safeguards of fable are blocking me on almost every task. I would like to see if fable is improved over opus for reverse engineering related work. Back to opus for me.
- ksimukka 4mo agoWow, credit to the safeguard team. I submitted my request about an hour ago to the cyber verification program and just now was approved.
- flessner 4mo agoI gave it a test spin. Half an hour and the 5 hour usage cap was hit in Claude Code. Not what I would expect on the Max 20x usage plan. I am sure it is great, but at this rate I would rather finish what I am doing with Claude Opus instead of structuring my usage around the 5 hour windows.
- synergy20 4mo agotruly scary. 2x at least token burning rate comparing to 4.8, can indeed run auto edit mode for hours. use it for super complex tasks then use cheaper model to do the rest, else will be broke.
- bogota 4mo ago[dead]
- bigboggerlogins 4mo ago[dead]
- sashank_1509 4mo agoI played textual chess with Fable. It took around 15 moves before it made a large blunder. I asked it to give its reasoning per move and it mistakenly assumed a piece was protected when it wasn’t and after the blunder it realized its fault and did not suggest an illegal move. Other LLMs lost game state far earlier. But a good human chess player can keep the game state in his mind much longer, so this random eval shows a big improvement over old AI models
- alleyio 4mo agohad an ancient, proprietary binary database format from the late 90s-early 2000s called 4d. opus 4.8 was great at figuring out how to extract the data, fable took it over the line with relative ease and completely reverse engineered the spec for 100% data recovery.
- weavoapp 4mo ago[flagged]
- asciii 4mo agojjj
- docstryder 4mo agoI've spent some time with Fable, and it is really good, definitely a step change from Opus 4.8, both for coding and general chat-style discussions. The vibes are incredible. There is an ease with which it solves problems and I've tested by replicating older chats in Fable - things that the older models found after 5-6 turns, Fable surfaces in the first response. It just gets things. Apart from all the above: the fact that they are intentionally writing this (that they degrade frontier LLM dev, silently vs loudly for biology/cybersecurity) in the system card is interesting to say the least - especially just before IPO. Notice that with this statement - that they're going to intentionally hobble the model for frontier LLM development - the general discussion has moved from, “Is the model actually that good?” to "they’re pulling the ladder up from behind them" That's actually super smart - wonder if Mythos (or the next unreleased model) had a say in coming up with that strategy (if it's intentional). Also - having access to extremely capable models before anyone else - which they have by default - is a incredibly advantageous position to be in.
- mrdependable 4mo agoHobbling the model may be smart tactically for them, but feels like it sets a really dangerous precedent.
- boombapoom 4mo agoits good for difficult problems, bad for design and code gen
- Georgecal 4mo ago[dead]
- jpcompartir 4mo agoAfter a day or so this is the first model that really feels next level compared to how Opus 4.5 felt on release
- jablongo 4mo agoQuestions about sentience and consciousness are being censored down to Opus 4.8 for me.
- Tyyps 4mo agoThe model is constantly switching to Opus for me, this is kinda unusable sadly.
- lellow 4mo ago[dead]
- staticman2 4mo agoFable is rejecting as unsafe analysis of poetry that uses formal medical anatomy terms. The guardrails are dumb as dirt.
- holysantamaria 4mo agoI am curious about this Fable 5 but maybe it’s just communication. I have been using DeepSeek v4 Pro to test it against Claude 4.6 and I couldn’t tell the difference… and it’s way, way cheaper. I don’t understand how American companies will survive the race. Maybe protectionism…
- stopyellingatme 4mo agoJust as an anecdote, i used it to review a PR with 24 file change. We pivoted from the initial draft to make a service bus subscription more lightweight and use SignalR to update the frontend. It used 1.4 million tokens and 34 sub agents during its review. This was not a large PR. So my read is that its very thorough, not good to use it for "small/medium" tasks unless precision is a very high priority.
- jgafni 4mo agoIt's only available temporarily, so I'm wary about falling in love with it or relying on it too heavily. Will it be part of a higher tier subscription?
- rambojohnson 4mo agopdf gives 404
- adithyaharish 4mo agoI found this error while using Fable 5 model in claude code. 400 api error. My advisior was on and it errored out saying claude opus 4.8 cannot be used as advisor while using Fable 5
- jackson281 4mo agoMythos 5 being unlocked for US gov cyber stuff is interesting. Wonder what kind of access other countries will get, if any.
- heugt 4mo ago[flagged]
- deleted 4mo ago[deleted]
- deleted 4mo ago[deleted]
- vitally3643 4mo agoAs per usual, the current Claude model's performance took a sharp nosedive the moment the new model was announced. Compared to the now-handicapped Sonnet model, Fable seems pretty smart I guess. But it also really, really wants to burn tokens. I asked it to look into a fairly straightforward database bug in my RN app, and while I was off getting coffee it decided to spin up an android emulator unprompted and started navigating the app by reading screenshots and injecting touch events. There went my entire week's tokens. There was no reason to even start the emulator, the bug wasn't graphical, so I have no clue what it was doing.
- greedydecode 4mo ago[flagged]
- jackson281 4mo agoThey claim it beat Pokemon FireRed with vision only, no maps or extra tools. That's cute but I'd rather see real-world benchmarks that matter, not games.
- jasonperez77 4mo agoMythos 5 being only for gov contractors feels like the old crypto wars all over again. Good AI for us, great AI for Uncle Sam.
- surcap526 4mo ago[dead]
- surcap526 4mo ago[dead]
- bicepjai 4mo agoIn Indian arranged marriages, families sometimes meet for an hour, everyone is on their best engineered behavior, and suddenly people are ready to make lifelong commitments based on smiles, tea, and a few photos. My mom would come home after one afternoon saying, “What wonderful people!” That is where we are with every new model release. The people yelling “Fable will take your job” are still at the first meeting. I have used it for 16 hours, and spent one of those hours fighting it over git. It rebased wrong, stashed changes, forgot it had stashed them, merged a stash it claimed did not exist, then reset to HEAD. By the end, I had lost the code we had just worked on. Maybe wait until first kid before minting trillionaires :)
- deleted 4mo ago[deleted]
- kunil4574 4mo ago[dead]
- arnarar 4mo agoI wish I had an opoortunity to use Fable 5
- norton2002 4mo ago[dead]