4 ms·
My understanding is they don't, not on their own: "We show that five SOTA LLMs (GPT-3.5, GPT-4, Gemini, Claude, and Llama2) struggle to recognize prompts provid
by dazilcher 3y ago
My understanding is they don't, not on their own: "We show that five SOTA LLMs (GPT-3.5, GPT-4, Gemini, Claude, and Llama2) struggle to recognize prompts provided in the form of ASCII art."
What the researchers are doing is include instructions for how to interpret ASCII art in the prompt, effectively "teaching" the LM how to parse it. Look at the article screenshots for an example.
Which means the ASCII art aspect is tangential, and the attack could be generalized using arbitrary encodings.
- vardump 3y agoGood catch. Somehow missed it when quickly skimming the article.
- phire 3y agoAh, that makes a lot of sense. One thing I learned playing around with chatgpt is that if you can convince the LLM to transform the input, that transformation happens internally, and the LLM can access this transformed version and operate on it even if it never outputs the post-transformed version. In fact, it gets worse. I discovered that if you can create an internal transformation, the LLM is now completely blind to the pre-transformed version. I only actually tried this with transforming data, but the LLM doesn't have any seperation between prompt and data, it makes sense it works for prompt too. So essentially this trick with ASCII art is a way to smuggle a prompt past any external prompt checker and only have it visible to the actual LLM. I guess the only "solution" is to use a version of the full LLM to do the content moderation, as it will be able to see the transform too. But that opens up other problems.