6 ms·
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down
by andy99 3mo ago
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
- afavour 3mo agoWhat are you asking that you’re so regularly running into censorship?
- icedrift 3mo agoIf you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in
- jefftk 3mo agoI thought we were talking about Opus 5, the model Fable now falls back to?
- eterm 3mo agoThis thread is full of people talking confidently about their experience with a model released just hours before. Either that or everyone is indeed talking across each other and talking about different things.
- bigbuppo 3mo agoI think the big take away is that Anthropic's products are hot garbage.
- wild_egg 3mo agoI'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.
- Retr0id 3mo agoI haven't been using it for long, but so far the refusals seem about on par with how things were on Opus 4.8.
- wewtyflakes 3mo agoI've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.
- msp26 3mo agoAsking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.
- skinfaxi 3mo agoWait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
- stavros 3mo agoAnything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.
- estearum 3mo agoProbably because the cost of blocking "is mitochondria the powerhouse of the cell" is nearly zero, while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite.
- areoform 3mo ago> while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case? Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?
- estearum 3mo agoDoes every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?
- thousand_nights 3mo agoi do homebrewing and asked it to compare some beer yeasts for me and hit the safeguards because... biology i guess lol
- dylanowen 3mo agoI was trying to debug/fix a segfault in the JVM which kept getting flagged
- arcanemachiner 3mo agoI was profiling a slow machine the other day, and triggered the safeguards. I've been saying this a lot lately, but it doesn't bites you until it bites you. The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
- cute_boi 3mo agoJust ask math question and it will censor that. Even Misanthrophic employee confirmed that.
- patcon 3mo agoWorking on dimensionl reduction algorithms, I hit it all the time. I'm also trying to port related protocols from single-cell transcriptomics to collective intelligence systems (working with people x reaction matrices as analogous to single-cells cell x gene matrices. Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me
- weird-eye-issue 3mo agoLiterally anything related to nutrition, athletic performance, etc especially if you ask it for research or sources
- AnotherGoodName 3mo agoWriting an implementation of a board game and one of the cards is called "microbes". Instantly knocked down to a lower tier model whenever it encounters that keyword because clearly bioweapons. Sigh.
- gck1 3mo agoReverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed. It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
- simpsond 3mo agoI’ve had codex/sol block me due to safeguards tripping without attempting to do anything nefarious. It’s not entirely predicable.
- maCDzP 3mo agoI have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.
- gck1 3mo agoThat's brilliant, I should try that. I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.
- alain_gilbert 3mo agoThe other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519... Fable understood it as something along the lines of: "introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED
- JumpCrisscross 3mo ago> Fable understood it as The dumbfuck bouncer Anthropic put in front of Fable decided this. Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.
- poopiokaka 3mo ago[dead]
- jpk 3mo ago> Fable is a PR model. Yeah, Fable is Anthropic's Cybertruck.
- JumpCrisscross 3mo agoNot quite. Fable is a Model S. The problem is you have to buy a Cybertruck to get it.
- spicybright 3mo agoBet it has something to do with that new model being blocked by the US government. It was blocked for like a month but now that it's finally released they put the safe guards waay up in fear of that happening again.
- theplumber 3mo ago
- Levitz 3mo agoI routinely get into blocks when running medicine-related material through it.
- jbritton 3mo agoI was flagged for basically answering yes to what Claude suggested to do, which was test commands on a port for my code for tests we had been discussing. I really think it was flagged simply because the words test and port were in the prompt. Their filters are pathetically poor.
- fluidcruft 3mo agoI haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.
- mdgld 3mo agoAre you more on the medicine side or the ML side? I don’t see many other self-admitted medical people on HN.
- kami23 3mo agoI've noticed a few self identify and other random occupations, cool to see those fields checking out this tech at a deeper surface level. I'm just a dev, but I appreciate the insights from other professionals.
- fluidcruft 3mo agoI'm more on the clinical physics side building tools for scanner/equipment QA and data handling/workflow automation/de-identification and anonymization but I also build random little tools to help optimize acquisition parameters.
- Tostino 3mo agoTrying to have it do some rework on a patch to Postgres I'm working on, it just completely shuts down. The reported issues were with privileged escalation and I was instructing it on how to fix.
- theplumber 3mo agoFable refuses to work on a login/signup system for example.
- idiotsecant 3mo agoThe better question is how would you not? I got demoted to opus from fable for asking if a cancer vaccine I saw on YouTube based on frog bacteria was a real thing. I've gotten it for asking how encryption works. It's incredibly touchy.
- d-m 3mo agoI took a photo of a rose bush and asked “what’s going on with this rose bush” which triggered a downgrade to Opus. It diagnosed it with rose rosette disease.
- areoform 3mo agoThings Fable's classifier has flagged, a non-exhaustive list, – "Does collagen supplementation empirically work?" - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?" – "Why do rabbits reproduce so frequently?" — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"
- alightsoul 3mo agoOh biology! It's dangerous!
- spicybright 3mo agoEven worse, it's offensive and kids could see it!!
- sigmoid10 3mo agoIf it goes on like this, American AI will eventually be able to solve the Riemann hypothesis but deny the existence of nipples.
- mdp2021 3mo agoFrom the late Perscheid: https://martin-perscheid.de/image/cartoon/3212.gif https://martin-perscheid.de/image/cartoon/3212.gif ("How do you reliably put an American out-of-combat.") But the same vibe is felt outside the USA anyway (the uk came close to that attitude with declarations from starmer earlier this year).
- lwansbrough 3mo ago[loads up most intelligent AI ever created] “Rabbit sex, how?”
- areoform 3mo ago
- rzk 3mo agoI had a long session about SQL with Fable and at some point, it started to falsely trigger censorship for any message I write in that conversation, even for the simple string: "random message."
- JumpCrisscross 3mo agoI wanted to explore some battery chemistry with Fable. It decided I was a terrorist, and then blew my session’s usage credits telling me to fuck off. The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)
- pmarreck 3mo agoWhat are you asking that you are NOT regularly running into censorship? Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer. Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.
- OkWing99 3mo agoAsk anything related to practical applications for quantum-computing, or space etc. The stuff you can find on Wikipedia. They were able to solve coding, but not what a real danger is.
- rdtsc 3mo agoNo op but my interactions with look like: Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible." Fable: "No. This is cybersecurity, blah blah, I won't help you" I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".
- actsasbuffoon 3mo agoI have way too many AI subscriptions. My favorite thing about Kimi K3 is that it just does what you tell it to do. “Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.
- actsasbuffoon 3mo agoI literally had Opus 5 flag a message because my hand was slightly too far to the side while typing. It apparently decided that a sentence that had some garbled words in it was a threat to national security.
- __MatrixMan__ 3mo agoI'm not who you're responding to, but I have a lot of questions about molecular mimicry: evolution pushes pathogens to be shaped like human cell surfaces because that way the immune system won't attack the pathogens (since, by doing so it would also attack the body). It's thought that many autoimmune disorders have an undiscovered pathogen as their cause, one whose mimicry caused such an attack. Discovery of these pathogens could be done computationally, I think. We can catch MHC binding event in process, find the bound protein, figure out which pathogens have genomes that code for proteins of similar shapes (epitopes), and we'd find--I hypothesize--a list of candidate pathogens for the cause of a delayed onset autoimmune disorder. Preventing these infections ahead of time would be a huge win against diseases like multiple sclerosis because without the initial exposure the immune system wouldn't have cloned so many of the cells that are attacking the host. Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.
- epistasis 3mo agoI use lots of biology in my day job. Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result. The memory aspect means that your prior work has a huge impact on what gets censored.
- smoe 3mo agoNot OP, but I got downgraded almost every time I had Fable implement something with security implications. Codex reviews it and identifies potential security issues, but Fable refuses to address them and downgrades instead.
- buzzerbetrayed 3mo agoYep. I cancelled my Claude Max subscription 2 weeks ago after feeling like Anthropic was doing everything it could to fuck with my day to day. Their lead would have to become significant for me to ever go back.
- pinkyboy 3mo ago[flagged]
- gck1 3mo agoEvery time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle. Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
- kccqzy 3mo agoJust go to /config. The very second configuration item is “Switch models when a message is flagged” and presumably you want to turn this off. Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?
- gck1 3mo ago> so you haven’t actually used Claude Code yet. Where do you think the principle came from? I've used claude code for a year, and stopped February this year.
- nananana9 3mo agoNot a great line of thought in general, sometimes the people who aren't doing the thing are the only ones worth listening to. "You aren't repeatedly slamming your head against the wall. Why would anyone listen to the opinion of a non-wall-head-slammer about the merits of wall-head-slamming?"
- kccqzy 2mo agoI don’t really fully trust people’s opinions unless they have first-hand experience. So I would trust someone who has tried the thing and found it to be bad idea, much more than someone who hasn’t tried it and is merely repeating common tropes or talking points.
- krzyk 3mo agoSo when you turn that off will you get answer from Fable?
- joinjune 3mo agoI told Claude Opus 5.0 to use a global api key for a PFAAS to deploy some web applications in a test environment and import some data into them. It balked at using a global api key because the security issues surrounding the permissiveness run afoul of it's sensibilities. I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues. I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues. Opus 5 is too cautious to be productive for me. It needs more tuning.
- NamlchakKhandro 3mo agoThe utter meme-think direction this company takes with regards to sycophancy of its models is disgusting.
- villish 3mo agoIt's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.
- mrloopex 3mo agoMaybe you just aren’t doing anything meaningful. Terrence Tao doesn’t whine about woke AI models.