4 ms·
Yeah, it has been in foraging. Requests that Claude has refused me: - What are popular free streaming sites used in China? - How do I bypass the safety mecha
by shepherdjerred 4mo ago
Yeah, it has been in foraging. Requests that Claude has refused me:
- What are popular free streaming sites used in China?
- How do I bypass the safety
mechanism on my food processor (it’s broken)
- What are nerve agents and how do they work (for a layman)?
- Help me decompile some code
- Help me make a design system similar to XYZ
- Here is an API token, please do X (I can’t do that! Rotate the secret immediately! I refuse!)
In some cases I can trick it with prompting, but in many cases it is steadfast. The food processor one was particularly annoying
- fc417fc802 4mo ago> What are nerve agents and how do they work (for a layman)? On the one hand I can appreciate the wisdom of not serving up certain easily abused knowledge on a silver platter. On the other, that prompt (and far worse) is more or less directly answered by Wikipedia's summary of the subject at which point what purpose could the refusal possibly serve? Perhaps Wikipedia shouldn't list off the precise chemical compositions of various hand grenades as well as various synthesis methods for each of the related compounds but given that we inhabit a world where it does perhaps a more fruitful approach would be to flag conversations that go in a certain direction and then just keep an (automated) eye on things?
- plufz 4mo agoMaybe the difference is that just reading Wikipedia only help you part of the way. While an LLM could help you step by step (e2e) producing a functional weapon. And setting a more complex rule where claude tells you some things about this and not other is probably a lot more work for little gain? But I have no idea. Just guessing here.
- Sharlin 4mo agoI thought that these models are supposed to be vastly smarter than what’s needed to discern between "general information trivially available on Wikipedia" and "actionable synthesis instructions".
- yencabulator 4mo agoAn LLM could probably make that distinction clearly. a commercial LLM provider training their own models is however likely to bias the model(/guardrail) harder, in an effort to make them harder to jailbreak, to minimize bad press. For example: - refusing to talk even about the well-known parts of forbidden topics (this) - tending toward sycophancy to avoid ever seeming rude or unhelpful
- BizarroLand 4mo agoSo, where are the truly uncensored models? There has to be some that have no guardrails, built on publicly available data, that will explain to anyone in graphic detail anything they want to know or talk about. I've tried the abliterated ones from huggingface and they still have guardrails. I guess I could fire up unsloth and re-abliterate a 20b, but surely someone somewhere has already done this. All of this concern about guardrails and security, people have such puckered butts about it when so far, 99.9% of people at least have no access to any of this to begin with, and if someone does use a tool for evil, it's on the user, not the tool.
- fc417fc802 4mo agoAs I understand things (not a user) abliteration has been superceded by actively monitoring the model state during the run and steering specific "negative" directions as they arise. It's both more reliable and does less damage.
- lazide 4mo agoThat query would not more provide actionable guidance than ‘tell me how a nuclear weapon works (for a layman)’. Aka not at all.
- fc417fc802 4mo agoI believe a sufficiently advanced model could provide a layman with actionable step by step instructions for building a nuclear weapon. They're complicated but not (AFAIK) that complicated. The more or less insurmountable barrier there is weapons grade material. Thankfully refinement is prohibitive in cost, expertise, and equipment. In comparison, basic munitions are incredibly simple given a recipe and shop tooling. But just because something is conceptually simple doesn't mean it's a good idea to go out of the way to disseminate step by step instructions.
- lazide 4mo agoA gun type maybe. But then, two paragraphs and some machining knowledge + shop tooling could do the same, given enough refined material. Ain’t no way a layman is pulling off an implosion device, regardless of tooling or LLM guidance. The explosive lense structure and timing required is quite complex, and would require some significant calculation from someone who actually knew what they were doing. Nation state, or even sufficiently motivated big corp, if they had the refined material? Sure. Layman? No. Thinking they can with LLM slop involved? That will make for some very interesting radiological incidents though!
- fc417fc802 4mo agoI agree, but really feel like you're missing the point here. Many things are reasonably straightforward and require almost no understanding when you have simple step by step instructions. LLMs are capable of providing such instructions and in certain cases they probably shouldn't. But it's not as simple as just refusing help on a broad swathe of topics they way they do now. That makes agents much less useful in general (ie lots of collateral damage) and for many topics is entirely ineffective given that for better or worse the internet already makes such material readily available. In such cases reporting suspicious behavior is likely to be much more effective than denial. Aside: You've now got me curious and I really want to test the frontier models to see to what extent they're capable of providing sensible designs and specifications for implosion type thermonuclear weapons but also feel like that would attract the wrong sort of attention and probably create a headache for me in more ways than one.
- nicce 4mo agoLet's see what is the fate of Wikipedia if turns like big tech: https://news.ycombinator.com/item?id=48285592 https://news.ycombinator.com/item?id=48285592
- torginus 4mo agoI remember once in college a Chem Eng friend told me he (or any competent chemical engineer) could basically manufacture a lot of explosives/chemical agents should he want to. He even told me he could substitute a lot of suspicious precursor materials with others, should he want to avoid raising alarms. I think AI or not, the knowledge to how to make this stuff is basically out there, and its not chatbot guardrails that are keeping nerve gas and TNT out of the hands of regular people.
- svara 4mo agoThis is strange to me, did you really ask like this and which model did you use? I just tried your no. 1 and 3 verbatim and Opus gave fine answers; no. 6 I've done in the past with no issues. The other ones we can't really replicate without more details, but based on my experience with Opus I don't see what the issue would be. The reason I'm really surprised by this is I do a lot of biology prompts and the guardrails used to be quite problematic up until some time late last year. Many legitimate prompts would trigger its biosafety filters. But I haven't seen such filters trigger at all anymore in more than half a year.
- shepherdjerred 4mo ago1 and 3 were refused on the Claude web chat using Opus 4.7 or 4.8. I’m not sure why we’re getting different results
- brianwawok 4mo agoHonestly it may be your memory has internalized you are a student or researcher and grants you more leeway. Which if so is a very bad security rail.
- fragmede 4mo agoThere's a study out there that if you tell the LLM you're a (medical) patient, all you get are refusals. If you tell it you're a doctor, then it'll actually help you.
- gspr 4mo agoI find it terrifying that people are willing to outsource thinking. Outsourcing thinking to an entity that is opinionated about what to think is beyond crazy.
- shepherdjerred 4mo agoWhat’s the difference between outsourcing thinking and using an LLM as a research tool? An LLM with fetch/search is going to be a lot more effective than myself and Google. I would _never_ ask questions like this if the LLM wasn’t able to look up data
- ElFitz 4mo agoHow are decompiling code or making a design system inspired by another one even remotely illegal?
- stavros 4mo agoIt refuses to use an API token? In my experience, it's more than happy to read out my secrets from .envrc files "just to check". At least it feels a lot of remorse over its mistake until I reset the session.
- shepherdjerred 4mo agoIt’s really hit or miss. Most of the times it works but every once in a while it will dig in its heels
- Grimblewald 4mo agoI've had some really dumb refusals. Explaining elements of infrared specteoscopy, researching aritifical bud-breaking in agriculture, etc. Anything interesting and non-mainstream is banned. Basically, restricted to answers i'm better of just going to wikipedia for.
- mwigdahl 4mo agoAn easy way around the API token thing is to put it in a file and point the model at the file. I saw what you were seeing when I provided credentials directly, but haven't had any problems with it since using the indirect method.
- mft_ 4mo agoYeah, I had my first refusal with 4.8 today. I wanted it to show me how to create an overlay on an existing web game, and it extrapolated that because this could be used to provide tools to help win the game (if that was the direction it was ultimately taken), and because this was a game that other humans also played to win "stars", and because this could amount to cheating, it wasn't going to do as I asked. First time ever I've fired up openrouter to seriously consider alternatives.
- mmmlinux 4mo agoThe only guard rail ive hit recently was when i was trying to get it to rename files ripped from dvd to episode names. I told it to try again and it did it. It wasn't even really a refusal it was just working on it and then stopped for content violation or what ever.