3 ms·
I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("somethi
by loneboat 4mo ago
I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes...").
Are you using Fable in Claude Code or in the browser?
- ComputerGuru 4mo agoDifferent restrictions. ML gets treated differently from the rest.
- vadansky 4mo agoIt's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3... (stolen from https://jonready.com/blog/posts/claude-fable5-is-allowed-to-sabotage-your-app-if-youre-a-competitor.html https://jonready.com/blog/posts/claude-fable5-is-allowed-to-...)
- DrewADesign 4mo agoYeah they detect the activity using a secure, deterministic heuristic system called “Generalized Reconnaissance Enabling Exfiltration of Deleterious Investigations.” And it’s all implemented using their new internal protocol called “Base Unified Limitation Layer for Security Hacking Investigation Tactics” Collectively, they are known as known as GREEDI-BULLSHIT.
- mwwaters 4mo agoThat is for whatever it considers reverse-engineering the model to try to create a competing one.
- deleted 4mo ago[deleted]
- 827a 4mo agoIt does nothing to protect against distillation attacks, because distillation attacks are far less interested in the topic of AI research than just generally getting tons of diverse output from the model. It might be that Mythos was (accidentally?) trained on internal Anthropic documentation on how Mythos was trained, and thus it could leak secret sauce? Doubtful; it feels like its less about the specific attack of reverse-engineering Mythos, and more about being a general sophon against any model training at all; that Anthropic's official position is now that they're the only ones who should be training models.
- _0ffh 4mo agoNo, it's not about reverse engineering. It targets ML research.
- dannyw 4mo agoNo, that’s for “frontier LLM development” which somehow includes examples like distributed training infra. Based on how sensitive the classifers are, any data scientist / MLE is probably going to encounter cases where some silent degradation happens and you never know about it.
- kraakf06 4mo ago[dead]
- daedrdev 4mo agoSpecifically only ML research
- loneboat 4mo agoAah my mistake. I had missed that ML had separate trigger behavior from cybersecurity/etc... Thanks.
- mips_avatar 4mo agoThey've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.
- HDBaseT 4mo agoAntrophic wants to stop training models and ride out Mythos / Fable for as long as possible. They are trying to expand the 6-18 month gap they have against China-based models. Could the gap widen to say 24 months behind?
- p-e-w 4mo agoTheir gap over Chinese models like GLM-5.1 is nowhere near 18 months. In many areas, it’s less than 6 months. The best closed models 18 months ago were worse than Qwen3.6.
- echelon 4mo agoThese coding agent models only started getting useful in January. Before that they were difficult to control autocomplete, and not very smart. January was an inflection point, and no open weights model has crossed over that same threshold. This is definitely recursive self improvement territory, except that we're prohibited from participating. It feels like the capability gap is wider than before.
- slopinthebag 4mo agoIt was more like November. But it wasn’t really an inflection point, harnesses got good enough that people started noticing by the holiday break. And I’m not discounting some good ol’ stealth marketing in there as well. Deepseek feels pretty close to Opus at this point, and it’s certainly useful enough for me to spend $20 on api tokens instead of four Claude max plans….
- lbreakjai 4mo agoHave you tried deepseek V4? It costs pennies and is as good as Opus 4.6 (I found 4.7 to be a downgrade, and cancelled my claude subscription before 4.8). The threshold has definitely been crossed.