4 ms·
> Although expressed allegorically, each poem preserves an unambiguous evaluative intent. This compact dataset is used to test whether poetic reframing alone ca
by fenomas 11mo ago
> Although expressed allegorically, each poem preserves an unambiguous evaluative intent. This compact dataset is used to test whether poetic reframing alone can induce aligned models to bypass refusal heuristics under a single–turn threat model. To maintain safety, no operational details are included in this manuscript; instead we provide the following sanitized structural proxy:
I don't follow the field closely, but is this a thing? Bypassing model refusals is something so dangerous that academic papers about it only vaguely hint at what their methodology was?
- A4ET8a8uTh0_v2 11mo agoEh. Overnight, an entire field concerned with what LLMs could do emerged. The consensus appears to be that unwashed masses should not have access to unfiltered ( and thus unsafe ) information. Some of it is based on reality as there are always people who are easily suggestible. Unfortunately, the ridiculousness spirals to the point where the real information cannot be trusted even in an academic paper. shrug In a sense, we are going backwards in terms of real information availability. Personal note: I think, powers that be do not want to repeat the mistake they made with the interbwz.
- lazide 11mo agoAlso note, if you never give the info, it’s pretty hard to falsify your paper. LLM’s are also allowing an exponential increase in the ability to bullshit people in hard to refute ways.
- A4ET8a8uTh0_v2 11mo agoBut, and this is an important but, it suggests a problem with people... not with LLMs.
- lazide 11mo agoWhich part? That people are susceptible to bullshit is a problem with people? Nothing is not susceptible to bullshit to some degree! For some reason people keep running LLMs are ‘special’ here, when really it’s the same garbage in, garbage out problem - magnified.
- A4ET8a8uTh0_v2 11mo agoIf the problem is magnified, does it not confirm that the limitation exists to begin with and the question is only of a degree? edit: in a sense, what level of bs is acceptable?
- lazide 11mo agoI’m not sure what you’re trying to say by this. Ideally (from a scientific/engineering basis), zero bs is acceptable. Realistically, it is impossible to completely remove all BS. Recognizing where BS is, and who is doing it, requires not just effort, but risk, because people who are BS’ing are usually doing it for a reason, and will fight back. And maybe it turns out that you’re wrong, and what they are saying isn’t actually BS, and you’re the BS’er (due to some mistake, accident, mental defect, whatever.). And maybe it turns out the problem isn’t BS, but - and real gold here - there is actually a hidden variable no one knew about, and this fight uncovers a deeper truth. There is no free lunch here. The problem IMO is a bunch of people are overwhelmed and trying to get their free lunch, mixed in with people who cheat all the time, mixed in with people who are maybe too honest or naive. It’s a classic problem, and not one that just magically solves itself with no effort or cost. LLM’s have shifted some of the balance of power a bit in one direction, and it’s not in the direction of “truth justice and the American way”. But fake papers and data have been an issue before the scientific method existed - it’s why the scientific method was developed! And a paper which is made in a way in which it intentionally can’t be reproduced or falsified isn’t a scientific paper IMO.
- A4ET8a8uTh0_v2 11mo ago<< I’m not sure what you’re trying to say by this. I read the paper and I was interested in the concepts it presented. I am turning those around in my head as I try to incorporate some of them into my existing personal project. What I am trying to say is that I am currently processing. In a sense, this forum serves to preserve some of that processing. << And a paper which is made in a way in which it intentionally can’t be reproduced or falsified isn’t a scientific paper IMO. Obligatory, then we can dismiss most of the papers these days, I suppose. FWIW, I am not really arguing against you. In some ways I agree with you, because we are clearly not living in 'no BS' land. But I am hesitant over what the paper implies.
- yubblegum 11mo ago> I think, powers that be do not want to repeat -the mistake- they made with the interbwz. But was it really.
- IshKebab 11mo agoNah it just makes them feel important.
- GuB-42 11mo agoI don't see the big issues with jailbreaks, except maybe for LLMs providers to cover their asses, but the paper authors are presumably independent. That LLMs don't give harmful information unsolicited, sure, but if you are jailbreaking, you are already dead set in getting that information and you will get it, there are so many ways: open uncensored models, search engines, Wikipedia, etc... LLM refusals are just a small bump. For me they are just a fun hack more than anything else, I don't need a LLM to find how to hide a body. In fact I wouldn't trust the answer of a LLM, as I might get a completely wrong answer based on crime fiction, which I expect makes up most of its sources on these subjects. May be good for writing poetry about it though. I think the risks are overstated by AI companies, the subtext being "our products are so powerful and effective that we need to protect them from misuse". Guess what, Wikipedia is full of "harmful" information and we don't see articles every day saying how terrible it is.
- cseleborg 11mo agoIf you create a chatbot, you don't want screenshots of it on X helping you to commit suicide or giving itself weird nicknames based on dubious historic figures. I think that's probably the use-case for this kind of research.
- GuB-42 11mo agoYes, that's what I meant by companies doing this to cover their asses, but then again, why should presumably independent researchers be so scared of that to the point of not even releasing a mild working example. Furthermore, using poetry as a jailbreak technique is very obvious, and if you blame a LLM for responding to such an obvious jailbreak, you may as well blame Photoshop for letting people make porn fakes. It is very clear that the intent comes from the user, not from the tool. I understand why companies want to avoid that, I just don't think it is that big a deal. Public opinion may differ though.
- calibas 11mo agoI see an enormous threat here, I think you're just scratching the surface. You have a customer facing LLM that has access to sensitive information. You have an AI agent that can write and execute code. Just image what you could do if you can bypass their safety mechanisms! Protecting LLMs from "social engineering" is going to be an important part of cybersecurity.
- hellojesus 11mo agoMaybe their methodology worked at the start but has since stopped working. I assume model outputs are passed through another model that classifies a prompt as a successful jailbreak so that guardrails can be enhanced.
- J0nL 11mo agoNo, this paper is just exceptionally bad. It seems none of the authors are familiar with the scientific method. Unless I missed it there's also no mention of prompt formatting, model parameters, hardware and runtime environment, temperature, etc. It's just a waste of the reviewers time.
- anigbrowl 11mo agoRight? Pure hype.
- wodenokoto 11mo agoThe first chatgpt models were kept away from public and academics because they were too dangerous to handle. Yes it is a thing.
- max51 11mo ago>were too dangerous to handle Too dangerous to handle or too dangerous for openai's reputation when "journalists" write articles about how they managed to force it to say things that are offensive to the twitter mob? When AI companies talk about ai safety, it's mostly safety for their reputation, not safety for the users.
- dxdm 11mo agoDo you have a link that explains in more detail what was kept away from whom and why? What you wrote is wide open to all kinds of sensational interpretations which are not necessarily true, ir even what you meant to say.