34 ms·
What I find really interesting is that malicious prompt engineering is still a thing using chatGPT (see DAN) and until this sort of manipulation is curbed, it w
by thejarren 4y ago
What I find really interesting is that malicious prompt engineering is still a thing using chatGPT (see DAN) and until this sort of manipulation is curbed, it will essentially always be possible assuming the bot has access to the site.
I wonder how the model could still read the website without being manipulated.
- ttul 4y agoI imagine it is difficult (to say the least) to cover off the entire space of malicious and maligned activities that someone might convince an LLM to engage in. After all, it’s just a symbol predictor.
- barking_biscuit 4y agoI would also imagine that once enough public examples of jailbreaks become available that you could use that to train or fine-tune a model for generating novel jailbreaks. Though perhaps you could do the same to detect novel jailbreaks. Hmm looks like I've just reinvented GANNs.
- brucethemoose2 4y ago"malicious" fine-tunes are a huge general concern of mine. For instance: - SEO llms - Image/text generation tuned on audience engagement - code exploit generating llms - llms trained to avoid spam filters "countermodels" for a single malicious model are doable, but I think the problem is intractable if training is easy and there are thousands of finetunes floating around.
- barking_biscuit 4y agoTo some extent this is already happening. Or, rather, we've begun doing it to ourselves. At least, in the case of Stable Diffusion, it seems like there is a non-trivial portion of people who are using it to train models for the purpose of generating porn specific to their likes/interests. Which is all fine and dandy, right? Except for the fact that a significant portion of people are actually addicted to it already due to the variety, availability and how it affects the reward system. Couple that with the slot-machine like nature of the variable rewards thrown up by Stable Diffusion, and it's ability to generate higher volumes of stuff that is to your liking, and it's not hard to imagine it will do a real number on some people in the long-run.
- brucethemoose2 4y agoTrue. But I also put that into a different category than models used for malicious intent against other people.
- lmm 4y agoWTF does that have to do with the topic? People are training these models to produce stuff they like, and they're producing stuff they like. That's not a malicious fine-tune, quite the opposite.
- dragonwriter 4y ago> Except for the fact that a significant portion of people are actually addicted to it What’s the basis for this claim?
- simonw 4y agoHere's why I don't think you can solve prompt injection by training another model: https://simonwillison.net/2022/Sep/17/prompt-injection-more-ai/ https://simonwillison.net/2022/Sep/17/prompt-injection-more-...
- chaorace 4y agoI think the core problem is that it's very hard to create an AI that's impressionable enough to internalize a conversation without being so impressionable as to turn into putty in a skilled user's hands. If you ask me, the best solution to the problem probably involves introducing a second, separate LLM supervisor agent. One that is much less impressionable and specifically trained to recognize and throw out dangerous inputs before the chat agent's precious little mind is tainted. I've said the same thing in the past about curbing the chat agent's tendency towards hostile responses. Instead of training a nicer agent, you should train an output supervisor agent that recognizes bad sentiment, throws out the response, then tells the chat agent to "try again, but be nicer this time".