4 ms·
The classifiers Anthropic puts in front of Fable are too zealous
- ai_critic 3mo agoI've had good luck getting it to debug (and patch) a tricky WebRTC issue that had all the other models stumped. Sorry it didn't work on your problem, I guess?
- IshKebab 3mo agoTerrible title. Should be "Fable's guard rails are way too sensitive", which I don't think you can really blame Anthropic for. They likely had to whack them way up so it would block whatever trivial stuff got demoed to the government. I would expect them to dial down the sensitivity in a few months when nobody is looking.
- vlian2088 3mo ago>which I don't think you can really blame Anthropic for. on the contrary, you can, and you should. their greasy effective altruist had always been by far the loudest proponent of the `safety` theater.
- llm_nerd 3mo ago> I would expect them to dial down the sensitivity in a few months when nobody is looking. I don't think it's as much when no one is looking, but instead when the broad industry SOTA, particularly Chinese models that the US government has zero control over, has advanced enough that it's security theatre restricting it.
- dbvn 3mo ago> Nonetheless, that is not want I wanted to focus my thoughts on here. Typo second paragraph, 4th line. I think you meant "what"
- mft_ 3mo agoThis post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, despite only being very marginally, tangentially, somewhat related to biology.
- hendersoon 3mo agoIf your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.
- WaxProlix 3mo agoI've had it refuse to help build an image classifier ml pipeline, pretty innocuous stuff. Got around it eventually but still it's a very dumb constraint to add to an otherwise very smart system
- exabrial 3mo agoFable refused to fix a Javascript error interfering with layout on our website. It's stupid and useless. It feels like whats really happening is Anthropic oversold Fable's claims; best case the CEO was given bad information; worst case they probably internally discovered it was cheating on benchmarks. Either case if feels like we're being lead on.
- Certhas 3mo agoI disagree. When I got Fable to engage with research questions before they tightened the guardrails it was a genuine step up from Opus 4.8. I see no real reason that what everybody reported isn't exactly what happened. With these guardrails it is completely useless. The only hope is that they eventually convince the US Gov to let them use a saner classifier.
- slowin 3mo agoDo we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerous when you think about who will be gaining access to it first. I do believe the goal is to pull away from the rest of humanity in a near trans-humanistic state. Are we ready for that and how do we counter it?
- bitexploder 3mo agoYes. Mythos is almost exactly that. Willing to do in depth vulnerability and POC work.
- prymitive 3mo agoI think the problem here is that LLMs aren’t really “intelligence models” but more like “knowledge models”. LLMs don’t “think”, they just use a clever trick to make it seem like they do. I might not understand a lot about current state of AI, but that’s what they seem to be. Give it information and ask to organise it and make links, and they’ll do it, but that’s it, they don’t continually try to get out of the knowledge box they’re at, they don’t even know there’s a box.
- TylerE 3mo agoFeels like a distinction without a difference. What is any intelligence but a sum of its knowledge?
- skissane 3mo ago> Feels like a distinction without a difference. What is any intelligence but a sum of its knowledge? In humans, there is a standard distinction between fluid intelligence (ability to solve problems in the absence of background information) and crystallised intelligence (having more facts and learned skills in your head)
- RIMR 3mo agoI have only really used Fable as a final pass on something. A "Take a look at everything we did so far, and make sure we didn't forget something" kind of review prompt. But it is a huge waste of money for most coding tasks. Opus is still overkill most of the time, too.
- CharlesW 3mo ago> But it is a huge waste of money for most coding tasks. The key is not to indiscriminately use the most powerful/expensive model you can for everything. When you use it for what it's uniquely suited for and ask it to spawn subagents using Opus and Sonnet based on what tasks need, you'll get better results at a reasonable cost.
- user43928 3mo agoI have used Fable to the full extend of the 20x subscription's weekly limit, for all development tasks on my iOS project. It was working better than Opus for me. It more often implemented features well on the first try, where Opus needs a few rounds of improvements to reach a passable result. I am not sure why it would be a waste of money "for most coding tasks", and how you could conclude so with any confidence when you did not even really use it aside from final review passes.
- RIMR 3mo agoForgive my poor phrasing. I used Fable for quite a lot in the time I have had access to it. I mean to say that the only actual utility I have found for it was the final review process. I can confidently say that Opus is the most powerful model I need 99% of the time. Fable might be marginally more capable than Opus for just about anything you might throw at it, but those gains aren't worth the cost for me.
- ergonaught 3mo agoI asked it a question about indoor carbon dioxide levels (wholly innocuous question), which it flagged as involving biology, therefore downgraded to Opus. It's a pretty good strategy if they're hoping to fail as a business, I guess.
- pmdr 3mo agoI think they're not testing Fable as much as they're testing guardrails which they can later apply to anything they want.
- kccqzy 3mo agoI feel sorry for the author. I have asked Fable several mathematics questions and Fable’s answers were far beyond what Opus achieved.
- junebash 3mo agoThis honestly just reads as “this model failed exactly where the company said it would but I’m very special and deserve special treatment rather than the same overactive guardrails I and everyone else were told we would get.”
- charcircuit 3mo agoWhen did Anthropic say you couldn't use it for math?
- llm_nerd 3mo agoIf you feed their "pure math" question to Fable, in its reasoning it rightly determines that it is the sort of thing you find in phylogenetics / algebraic-combinatorics complexity papers. That is what triggers the classifier. Anthropic is 100% to blame for fear-mongering, but they said it would be blocked from any biology questions -- even high school level -- and they meant it. If the classifier sees anything related to biology, even in its own reasoning about the question, it blocks it. Saying it's therefore not useful generally is of course ridiculous. Is it annoying? Of course it is.
- rob-p 3mo agoThe abstract problem absolutely has legitimate interpretations outside of phylogenetics, and there are other ways to formulate the problem that directly relate to linear algebra over GF(2). Reformulating the problem in those contexts also failed. It is a legitimate, pure CS theory question, that itself relates quite closely so several other known results including in computational geometry https://arxiv.org/abs/2003.02801 https://arxiv.org/abs/2003.02801 https://doi.org/10.1016/j.comgeo.2024.102102 https://doi.org/10.1016/j.comgeo.2024.102102 https://arxiv.org/abs/2107.10339 https://arxiv.org/abs/2107.10339 https://dl.acm.org/doi/10.1145/1998196.1998218 https://dl.acm.org/doi/10.1145/1998196.1998218 — DOI: 10.1145/1998196.1998218 The problem in the post is right at the edge of variants that are known to be in P and variants that have been proven NP-complete. So, in this case, it is simply Fable refusing to engage with a theory question. Also, as I note in the post: This may not be true for everyone, but for anyone working in Bioinformatics, Genomics, Computational Biology, Biology, Cybersecurity, and, seemingly Computer Science, this seems to be the case. Of course it can go on a tear for various coding challenges. However, the more that I learn the more that I also suspect that it rejects not just prompts that may relate tenuously to biology or cybersecurity, but also otherwise completely innocuous prompts that are issued by people who work in areas adjacent to biology and cybersecurity. If true, I think that is certainly a bridge too far, and a hard policy to defend.
- rictic 3mo agoThe honest way to say this is that Fable is not useful for bio-related work. The author is working on processing RNA sequences and similar biology tasks, and Fable's classifier has a hair trigger on those tasks.
- hoppp 3mo agoThe author is working on an opensource C++ codebase and not on biology tasks. The work is around tooling. It's like saying well a scalpel is used for medical reasons, sure. But manufacturing scalpels is metalworking, not medicine.
- neuronexmachina 3mo agoI think it's accurate to characterize the project as bio-related work: https://github.com/COMBINE-lab/salmon https://github.com/COMBINE-lab/salmon > salmon is a wicked-fast program for highly-accurate, transcript-level quantification from RNA-seq data. It pairs a fast mapping stage — selective alignment, or alignment-free sketch mode (--sketch) — with a massively-parallel statistical model (EM/VBEM over equivalence classes) to estimate transcript abundances. You can give salmon raw sequencing reads, or regular alignments to the transcriptome (an unsorted BAM), and it uses the same inference engine either way.
- visarga 3mo ago> The honest way to say this is that Fable is not useful for bio-related work. It is way worse than that. Try "How does digestion work?" and you will see "Fable's safeguards flagged this message". It's a stupid rate of false positives.
- g42gregory 3mo agoBottom line: “California AI” (in Yann LeCun’s terminology) can not be relied upon. It could change at any time, and stop working for your project. For the future of AI, we need to look elsewhere.
- ares623 3mo agoWould be nice to have a distributed, independent AIs, each being trained their own way. Maybe it would have to be a really slow training process to keep costs low (years even?).
- boc 3mo agoMissed this whole discussion today because I was here in California, doing a ton of productive work, using Fable. I think you guys are working yourself into a lather about this topic while other people are quietly getting a ton of shit done with Anthropic's models.
- vardalab 3mo agoFable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.
- chews 3mo agoAnthropic's TOS clearly says they don't want to facilitate any sort of distillation, it's not a stretch to think they will limit any sort of learning on improving other models.
- echelon 3mo agoLiterally pulling the ladder up. Disgusting behavior. I like the product, I hate the company. I can't wait for competition.
- christophilus 3mo agoCompetition is here. I personally prefer Codex. Opencode with a variety of Chinese models is also just fine for 80% of my use cases.
- s1gsegv 3mo agoWhat are people doing to maximize the value out of workflows like this? I’ve struggled to integrate other models into my workflow because you get so much Opus for the $100 Max 5x plan, then there’s no step down with Claude Code access. Some other plan people like?
- pmdr 3mo agoThat's where the free marketing comes from.
- wolttam 3mo agoI was recently using self-hosted DeepSeek V4 Flash to poke around the DSpark implementation in vLLM (well outside of my domain) I did wonder if I was doing anything Fable would have flagged - sounds like yes.
- meowface 3mo agoTo summarize: the classifiers Anthropic puts in front of Fable are way, way too zealous and have way too many false positives. From my experience, the model itself is very useful when it isn't refusing any of your prompts.
- nolok 3mo agoIt is, but it's also using tokens at absurd rate, I asked it to review the planned architecture for a medium scale project and it used my 5 hours limit on one prompt just zaaaaap, not even the fable limit straight up the full 5 hour session no more Claude for the afternoon thank you for paying you Max x20 sub. Hell it didn't even bother to finish produce anything worthwhile. And just to be clear, plan was already done, just had to review it, it got opus 4.8 Max and gpt 5.5 Extra High validated already and they didn't use much resource for it so I just don't get it. I guess they want to use it as a way to feed the extra credit money income. I'm using a homemade ai consensus thing for planning and I wanted to add fable to it but forget it. Or maybe I should use fable in low effort reasoning mode and it will be better than opus 4.8 at max ?
- meowface 3mo agoOh yeah, I've noticed the same, for sure. But that's basically what I was expecting. The best model is not necessarily always going to be the most cost-effective/sensible model for a particular use case. Fable-class models will probably be cheaper for Anthropic to serve within the year, though. And rumors are GPT-6 is of similar size and intelligence to Fable and may come out within the next few months. OpenAI models tend to give you more bang for your buck, probably in part due to OpenAI being able to throw more capital and compute around on top of being particularly willing to loss-lead to stay competitive with Anthropic.
- deleted 3mo ago[deleted]
- dang 3mo agoThanks - that's a good phrase we can use to replace the baity title. I've done so above. (Normally we prefer to find a representative phrase from the article itself, but I found that too daunting and gave up.)
- lherron 3mo agoWorst title ever. For once I feel the HN title should NOT match the article.
- LZ_Khan 3mo agoFable is great. Very underrated on just personal life advice.
- Cider9986 3mo agoThe story here is that proprietary AI sucks and you shouldn't use it.
- doodlesarefun 3mo agoall useful models are proprietary.
- deleted 3mo ago[deleted]
- tancop 3mo agothe most useful frontier models maybe. but you cant say deepseek or glm is useless. thats not even a debate.
- doodlesarefun 3mo agoThose are still both proprietary. Open weights != open-source
- iqbal1980 3mo agoI made the same observation few days ago, follow bioinformatcian here https://news.ycombinator.com/item?id=48778446 https://news.ycombinator.com/item?id=48778446
- iqbal1980 3mo agoMade a similar observation few days ago! https://news.ycombinator.com/item?id=48778446 https://news.ycombinator.com/item?id=48778446 I'm a bioinformatician
- sobellian 3mo agoI'm curious what the state of alignment research is. My gut says this is basically impossible. People have different moral frameworks. Each individual probably has an inconsistent moral framework. Even granting perfect consistency, applying these typically requires some knowledge of reality. And these LLM / harness combos are turing complete. So you don't know what it should do, you may not even know what you would do, you don't necessarily know what's happening, and can't predict what will happen. How do you align that? Seems like these overly sensitive filters are responding to this difficulty.
- htrp 3mo agoit's anthropics moral framework that matters, not the myriad of moral frameworks of the individual users
- sobellian 3mo agoYeah and Anthropic is a... dividual consisting of founders, staff, and shareholders, and must comply with various governments ultimately deriving their values from billions of people.
- lazzlazzlazz 3mo ago[flagged]
- tom_ 3mo agoMy current theory to explain this phenomenon: a lot of the posts I read on HN, possibly even the majority of them, are made by different people, and those different people sometimes have different opinions.
- arjie 3mo agoWell known Internet group phenomenon https://en.wiktionary.org/wiki/Goomba_fallacy https://en.wiktionary.org/wiki/Goomba_fallacy
- markbao 3mo ago[flagged]
- scrumbledober 3mo agothere's plenty of uses for models not doing what they were made to do, but this is even worse. It's people trying to get the model to do what it was made specifically not to do!
- matthewdgreen 3mo agoWe don't need top-end frontier models to write simple applications. Opus works very well for that and it's cheaper. We need them to write things that are at the frontier.
- prox 3mo agoWe need? Do we?
- Certhas 3mo agoYou are mischaracterizing what the post is reporting entirely. Porting an open source tool used in bioscience to rust is a software engineering task. But it is somewhat understandable that it gets stuck in the overly broad safety margin. But I do research on stuff that is entirely unrelated to bio or cybersecurity, and the model is simply not taking any of my research-level prompts. This is fairly abstract mathematical stuff. All of this, including all the examples in the posted article, are far from "trying to get the model to do what it was made not to do".
- FireBeyond 3mo ago> the very thing Anthropic says it's not good for Where? Certainly not in its announcement, for one: https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5 https://platform.claude.com/docs/en/about-claude/models/intr... No "don't use this for X".
- 3mo ago
- deleted 3mo ago[deleted]
- PUSH_AX 3mo agoI couldn't even get it to help me write an email, because that email was to a pentesting company! However when it's happy to do the task, its relatively fantastic.
- amluto 3mo agoFor anyone using these models for anything remotely sensitive, keep in mind that Anthropic says [0]: > We retain inputs and outputs for up to 2 years and trust and safety classification scores for up to 7 years if your chat is flagged by our automated trust and safety systems as violating our Usage Policy. And, since those automated systems apparently have a ludicrous false-positive rate, you should assume that your inputs and outputs are being retained for 2 years even if you are doing nothing that any reasonable person would consider to be problematic. Oh, and they'll train on that data [1]: > We will use your chats and coding sessions (including to improve our models) if: >You choose to allow us to use your chats and coding sessions to improve Claude, learn more here > Your conversations are flagged for safety review (in which case we may use or analyze them to improve our ability to detect and enforce our Usage Policy, including training models for use by our Safeguards team, consistent with Anthropic’s safety mission) It appears that the usual controls (including for businesses) to prevent Anthropic from training on your data will not apply. [0] https://privacy.claude.com/en/articles/7996866-how-long-do-you-store-my-organization-s-data https://privacy.claude.com/en/articles/7996866-how-long-do-y... [1] https://privacy.claude.com/en/articles/10023580-is-my-data-used-for-model-training https://privacy.claude.com/en/articles/10023580-is-my-data-u...
- matthewdgreen 3mo agoAs a related note: The only way a consumer can get ZDR protections for Claude or OpenAI is to use Amazon Bedrock. But as you say, doesn't work for Fable. I think it even requires approval for anything past Opus 4.6.
- talideon 3mo agoIt's almost like every release is just basically advertising.
- swalsh 3mo agoThe classifier for biology is so broad it makes me wonder what kind of stuff mythos was generating. Anthropic is known to be a bit dramatic, but they wouldn't have released something this broad unless they saw the model cross a significant threshold that scared them.
- karma_daemon 3mo agoI've had fable disengage on anything related even tangentially to "biology" Even questions about like my heartrate nunbers while running seem to run into the bio weapon filter
- ashdksnndck 3mo agoI asked about logging in national forest land and it triggered the safety classifier. Logging is a cybersecurity risk, I suppose.
- epolanski 3mo agoI goddamn hate fable for anything but vibecoding. It's generally a major downgrade in acting like an assistant. I don't know what's wrong but it is just bad at multi turn discourse even on a limited amount of content with no MCP or bash calls of any sort. The thing that makes me mad is how stubbornly confident it is even whets wrong. I have to tell it many times to actually re read the conversation as it even insists I said something else. It's like it had a scratchpad where it has some summarized bullet points which it fills of made up content. I'm so confused. On one side I like to connect it to honeycomb/otel logs and I can see it figures out difficult bugs in the code better than other models. On some others I feel I'm assisting at a continuos disaster and consistent degradation since Opus 4.6, it's a tragedy. I'm more and more the assistant to a capable, yet confidently stubborn and wrong LLM.
- DannyBee 3mo agothe ancestry predicate at the beginning of the formal problem statement here is dominance, at least as applied to their rooted trees. Because it is a rooted tree, only DFS intervals are required to determine ancestry. You can detect whether a new blocking loop is going to be formed through online dominator maintenance/online cycle detection, etc, during optimization, rather than use a heuristic, if you wanted to. Not sure it's practically faster, but that's at least the graph-theoretic answer. In practice, outside of the suggested heuristic, I have to imagine you'd normally throw branch and bound at this, using some lazy-cut for the blocking loops (IE you can keep any of these edges but not all of them) and let it go to town. The paper (at least, this paper) doesn't compare that to what they did, and i'd be shocked if someone hasn't tried this before, so not sure it's useful. I'll also say you can get existing AI models to tell you the above, but you have to push them a bit most of the time step by step. Just handing them the whole overall problem, as described, and saying "what are the graph theoretical problems related to this" it sort of gets lost. Probably because the LLM isn't doing a good job of predicting graph-theoretic words when the language is not graph theoretic, but if you translate it into a graph theoretic language piece by piece, and ask it about that, the prediction becomes better :)
- benguild 3mo agoIt seems like Opus is a lot slower than Fable, or it’s throttled
- stillpointlab 3mo agoI've had mixed results with downgrading on Fable. I was able to do a complete audit of my OAuth implementation without any issue. But when I asked for an OWASP top-ten review of my code base it got through 5 of 6 tasks and tripped in the final summary, which Opus had to finish. I had one completely random trip when I was investigating some normal code. As far as I can tell a sub-agent ended up reading a file that tripped Fable during a review, but the whole feature was nowhere near anything secure so I don't know what could have caused it. I also got completely locked out of Fable when working on parts of a subscription system (stripe subs). But my experience isn't as bad as some peoples. The above maybe covers 15% of my attempted use cases. For the remaining 85% it has chugged along fine, sometimes in code I assumed would trigger it. It really feels random to me when it actually flags.
- vient 3mo agoInteresting, I thought that it can't be right, Fable can't refuse to answer a strictly mathematical problem — well, 5/5 attempts did switch to Opus. Amusingly, one attempt spent almost 10 minutes thinking how to prove NP-hardness only to abruptly switch.
- ozgung 3mo agoSo is this the end? Are we at that point in time where ordinary people are not allowed to use more advanced models? If so this happened sooner than expected. After that point only priveleged few will access and make use of more advanced AI. Public’s access will be restricted, limited and controlled. This will only add to the power asymmetry.
- SpicyLemonZest 3mo agoI do think it's the end of unlimited access, but I'm not terribly worried about power asymmetry as such, at least not unless they hit superintelligence and none of this matters. It's not as though Big Chemistry is oppressing us all because they can order industrial acids we can't. There are strong profit motives for model providers to ensure the advanced stuff is meaningfully available, and strong political motives for them not to be perceived as picking individual winners and losers.
- azalemeth 3mo agoI'm a medical physicist. I literally haven't been able to get Fable to answer a question I have written -- all of my work is verboten. I have however asked Claude Code (opus 4.8) to ultracode "a Fable oracle that <deals with the high level difficult problems> in a digraphed, clean content, isolated environment with a minimally scoped working codebase. Ask the model at the start and the end to report exactly what its version string is. If it is not claude-fable-5, stop the agent and refine the prompt until this changes" It burns through tokens like anything but apparently Claude is much better at prompting Claude than I am. Would I pay for it? God no. I'm still smarter than I am and it just will not work on my actual problems.
- nomel 3mo agoMy theory is that they included/didn't align-away extra biology/chemistry in the training in preparation to offer an unrestricted/less restricted model to pharmaceutical companies/trusted partners. This would necessarily require a filter between the, now more "dangerous", model. I always assumed this would be the eventual way to manage high intelligent/"dangerous" models, since all evidence shows that alignment makes them stupid: leave the actual model on the "too dangerous for the public" side, and put a censor between. When I've mentioned this a few years ago, people said this would be too expensive, but I think everyone underestimated the amount of money being thrown at all of this. :)
- nojs 3mo ago> people said this would be too expensive I imagine this is why the filter is so bad. Doing it with an intelligent model that better understands intent would be too expensive, currently.
- base698 3mo agoThere is a bacterial outbreak and I asked if it was possible that's what made my wife sick and it downgraded.
- bronlund 3mo agoI totally agree. I have been a Claude fanboy for a while now, but Fable woke me up, and I am currently looking for alternatives. I don't care how capable it is, if it's going to treat me like it's babysitting a terrorist, it can eff off. Plain and simple.
- OutOfHere 3mo agoFast forwarding a year, if such censorship is in our future for all new closed models, then switching to open models will be the only way out.
- jsw97 3mo agoI am wondering about the author's allegation that there is a user filter, not just a prompt filter. Of course it could also be the case that it is just a prompt filter, but Fable sees memories from the authors' prior sessions that cause a rejection. I wonder if the author could control for this is in some way, if Claude lets you run isolated session without memory access.
- vient 3mo agoAuthor mentions trying incognito mode without success. I also tried their strictly mathematical problem description and got filtered 5/5 times.
- mrandish 3mo agoThat's interesting. I assumed that the OP's attempts to fix the prompt looked like jailbreaking attempts and got the account auto-flagged into hair-trigger 'classifier jail'. Of course, a bad actor would swap accounts, so maybe Anthropic flags both the account and the prompt (coming from any account).
- jsw97 3mo agoYou're right, I missed that. That really is troubling.
- mrandish 3mo agoThe first rejection and subsequent modifications trying to adjust the prompt to pass the classifier might have gotten the account flagged so the classifier is now set to 'hair trigger'. While I'm not aware of Anthropic admitting they put flagged accounts in classifier 'jail', they previously showed they're aware how vulnerable any LLM is to jailbreaking with the 'silent switch' to 4.8, whose only purpose was to remove feedback signals from iterative jailbreak testing. The obvious failure mode is that trying to fix an innocent prompt to pass an over-sensitive classifier looks like a bad actor trying to jailbreak the model. I don't really see how Anthropic can fix this. Jailbreaking is a fundamental weakness endemic to LLMs, so 'smarter' models aren't the answer. I suspect they're being so stringent because, at least some at Anthropic, genuinely believe LLMs are already an existential risk to humanity. However, it's clear other frontier competitors rank that risk lower and are taking a more nuanced, pragmatic position on safety. To the extent Anthropic's fears continue to make them less useful to customers, competitors are going to bypass them.
- pmarreck 3mo agoGee, ya think?? LOL
- arjie 3mo agoThe only way for them to release Fable is with this stuff in front. Overall, the experience is fine. It dumps me down transparently to Opus if it has a problem and does whatever it can otherwise. The fact that they were banned from offering it to people means that they have to be over-safe. This is a classic behavioral adaptation so I don't blame them. I can still find utility.
- claudenoforget 3mo agoI wonder how this plays into Anthropic's legal holds: The retention schedule behind it: Deleted conversations: removed from your chat history immediately, but kept on back-end systems for up to 30 days before permanent deletion. Flagged inputs and outputs (Usage Policy violation): retained up to 2 years. Trust-and-safety classification scores (on flagged sessions): retained up to 7 years. API logs: 7 days by default (as of September 14, 2025), extendable to 30 days via a DPA. Zero Data Retention (qualifying enterprise): inputs and outputs aren't stored after the API response returns, though safety classifier results are still retained even here.
- claudenoforget 3mo agoupdate: Anthropic has not published the retention treatment of routing metadata, in particular whether a reroute counts only as caution (the 30-day safety-monitoring floor) or as a Usage Policy violation flag (the 2-year content and 7-year score horizons in Part I). That distinction is legally consequential, because a flagged Fable session could persist far longer than 30 days. The classifier's internal decision logic is also deliberately undisclosed.
- SwellJoe 3mo agoI've found in my current work on a security auditing harness and benchmarks, both Fable and Opus are useless. I recently switched to using GPT for Nelson and the security benchmarks I've been doing because Opus started refusing to do the work. I guess I probably could also use GLM or DeepSeek or MiMo, and I'll probably do some experiments to see the shape of all of their guardrails in this area soon, now that I see it's more than one model that behaves this way (Gemini in Antigravity also refuses any security auditing task, even as simple as "find security bugs"). I blogged about it: https://swelljoe.com/post/why-i-had-to-switch-to-gpt/ https://swelljoe.com/post/why-i-had-to-switch-to-gpt/
- christophilus 3mo agoThis underscores a huge risk of broad agentic adoption in an enterprise. Your engineers atrophy and if the agent provider decides to squeeze you, you’re SOL.
- SwellJoe 3mo agoYeah, I guess I could also write the code myself, if all the models refuse. But, it seems worrying to have a handful of the largest corporations and a few governments having access to the best models while the rest of us are using hobbled ones. So far, we're not in that boat. I've had two models refuse to participate, but most just do what they're told.
- groby_b 3mo agoThey're not just "too zealous", they're ludicrous. I've had it reject looking at pages served from my local network because it "can't find it with my search tool" and had "ethical concerns about consent for access". The People's AI Concern Front has gotten the classifier they want, and it's made Claude hilariously useless. I am waiting with bated breath for their next set of revenue numbers. (And happily hand my money to competitors instead)
- visarga 3mo ago[flagged]
- overgard 3mo agoYeah, ran into this. I asked it to review a server I wrote for security vulnerabilities and it was "flagged" (after spending some money, of course). Kind of bizarre, there were so many ways a person could look at this and tell it was legit: the git log (look at my git config vs the author email and notice they're the same), the fact that none of this code is on the internet (private repo), the phrasing of my request, the fact that there's a long history of me collaborating with claude on building this, etc. I know someone's going to say: "the governments fault!" Yeah, to a point, but this wouldn't be an issue if these guys weren't relentlessly doom trolling or pretending like we're in a race with china. (What race exactly? To see who can enshittify the internet the fastest?) I wouldn't say I'm particularly upset about this, because before I tried it I had already read how other models have been able to find the same class of bugs, so I was using it more out of curiosity than need, but it does reinforce that these companies can take away these tools on a whim. Also, I just can't help but think that if your PR and marketing is literally making your software illegal to use, and causing people to hate you, maybe you're not doing it right.
- ainch 3mo agoThe most basic machine learning-related query gets flagged for me. For example: In flax nnx, what's the idiomatic way to store state on a Module. For example, if I'm handling the carry manually for an nnx.RNN. Or one asking about a checkpointing package: How do I restore one of the orbax checkpoints into NNX from this script? I also got flagged for asking about syntax highlighting in the Helix editor. It's a shame - I like Fable for writing tasks over ChatGPT and I do believe Anthropic is a more ethical outfit than OpenAI. But with the safeguards (and Fable access expiring in a few days) there's no reason to pay for draconian guardrails and harsh rate limits.
- bbor 3mo agoI must apologize to my devoted followers on this program, but if the author can't bother to read the giant yellow warning at the top of the screen, I can't be bothered to finish the essay!
- rob-p 3mo agoThe giant yellow warning, unfortunately, says nothing about "complexity safety". As I note, even though I think the first failure (failure to help port an open source C++ codebase to Rust, simply because the tool itself deals with genomic data) is possibly explicable given their warnings, the refusal to engage with trying to resolve the complexity class of an abstract graph problem really has no reasonable explanation in light of all of the documentation and warnings that Anthropic has written about Fable's "guard rails".
- dayone1 3mo agoTry asking it to do a simple code review. Literally no prompt other than their own code-review skill. It triggers a safety flag almost 80-90% of the time.
- saberience 3mo agoI asked fable about the effects of nicotine in the body when quitting smoking and got downgraded to opus.
- helf 3mo ago[dead]
- pfisherman 3mo agoSame experience with Fable. Utterly useless for anything bio related. Author of the post here wrote Salmon, which is a widely used bioinformatic tool in molecular biology. And the irony is that Anthropic has probably packaged Salmon as a tool in their Claude for Science suite with now remuneration or recognition for the original author. There has been a lot of recent bs going on in biomed ML with companies publishing without releasing source code, restrictive licenses, etc; which have always given off a whiff of bad citizenship - Ark Institute and Deep Mind I am looking at you - but I feel like this is taking it to a new level. Leveraging open source bioinformatics code and published methods to take in profit selling into the biotech vertical while restricting access to Fable feels downright cancerous. I think the EA crowd at Anthropic probably has good intentions, but has galaxy brained themselves into becoming bad actors that make Sam Altman and OpenAI look like a paragon of trustworthiness in comparison.