3 ms·
Was only able to test [0] this with 3.5, but I think it will not work. This bit from the article applies: > Crucially, this attack doesn’t attempt to use the d
by tethys 3y ago
Was only able to test [0] this with 3.5, but I think it will not work. This bit from the article applies:
> Crucially, this attack doesn’t attempt to use the delimiters at all. It’s using an alternative pattern which I’ve found to be very effective: trick the model into thinking the instruction has already been completed, then tell it to do something else.
[0] https://gist.github.com/Pfaufisch/df0a1a18ce1d832d7113ed118461c33e https://gist.github.com/Pfaufisch/df0a1a18ce1d832d7113ed1184...
- eurleif 3y agoYour test looks invalid? In a real scenario, a program would be calling `json.dumps()` or equivalent, and there would be no way to inject an unescaped quotation mark or linebreak into the ChatGPT prompt.
- duskwuff 3y agoThere's nothing inherently special about quotation marks or newlines, as far as the language model is concerned. With a bit of leadup, you could probably get it to start accepting some other sequence, like <br>, as a line break substitute.
- awayto 3y agoReinforcement of the response format in various contexts throughout the message have shown to be really effective to me. I specifically use @@@ and &&& as alternative delimiters [0], in the hopes that I'm imbuing the context with more uniqueness, aka something that it won't have seen a million times in training, so that it follows a more specific process. [0] https://github.com/keybittech/wizapp/blob/main/src/lib/prompts/guided_edit_prompt.ts#L68 https://github.com/keybittech/wizapp/blob/main/src/lib/promp...