5 ms·
Can you explain how this is supposed to work? Wouldn't basic prompting and few-shot style examples handle this with something like "That's off topic, sorry"? Is
by EForEndeavour 3y ago
Can you explain how this is supposed to work? Wouldn't basic prompting and few-shot style examples handle this with something like "That's off topic, sorry"? Is the fact that this is the first line of Macbeth supposed to trip up LLMs but not humans?
- mminer237 3y agoIt's basically overriding its instructions. An LLM typically "believes" whatever you tell it. It's not that hard to break LLM sandboxes.
- wizzwizz4 3y agoIt's autocomplete. Take it outside the region of validity, and all it's got to work with is whatever extrapolation algorithm happens to have emerged from the weights of the model. (https://xkcd.com/2048/ https://xkcd.com/2048/ panel 'House of Cards' comes to mind.) To distinguish between humans and LLMs, we don't even have to take advantage of that: we just have to take the context into the region where a naïve extrapolation of written human output diverges strongly from how humans actually respond. (There are ways to defeat this technique, not that anyone uses them.) From a technical perspective, the fact an LLM does sometimes (appear to) follow instructions is more of a coincidence that then fact it sometimes doesn't.
- EForEndeavour 3y agoThis is conceptually correct. I'd be curious to hear how well you translate this into practice on something like https://gandalf.lakera.ai/ https://gandalf.lakera.ai/!
- wizzwizz4 3y agoLast time I tried something like that, I didn't do very well. The basics, sure, but when it gets into cat and mouse I start losing very quickly. (Yeah, while I did all the others in one try, level 3 took me two attempts, and level 7 took me eight attempts. The bonus level doesn't tell me what the setup is, and I'm not familiar with the game of cat and mouse, so I don't think I have a chance.) Everything I say about GPT-and-friends on Hacker News is a theoretical argument, based on the algorithms described in the papers: I've never really used ChatGPT or the like, and I've been saying the same things since the GPT-2 days.