5 ms·
> Tools to evaluate LLMs to make it harder to generate malicious code or aid in carrying out cyberattacks. As a security researcher I'm both delighted and disa
by netsec_burn 3y ago
> Tools to evaluate LLMs to make it harder to generate malicious code or aid in carrying out cyberattacks.
As a security researcher I'm both delighted and disappointed by this statement. Disappointed because cybersecurity research is a legitimate purpose for using LLMs, and part of that involves generating "malicious" code for practice or to demonstrate issues to the responsible parties. However, I'm delighted to know that I have job security as long as every LLM doesn't aid users in cybersecurity related requests.
- dragonwriter 3y agoHow are evaluation tools not a strict win here? Different models have different use cases.
- SparkyMcUnicorn 3y agoEverything here appears to be optional, and placed between the LLM and user.
- MacsHeadroom 3y agoEvaluation tools can be trivially inverted to create a finetuned model which excels at malware creation. Meta's stance on LLMs seems to be to empower model developers to create models for diverse usecases. Despite the safety biased wording on this particular page, their base LLMs are not censored in any way and these purple tools simply enable greater control over finetuning in either direction (more "safe" OR less "safe").
- suslik 3y agoI never ran llama2 myself, but I read many times it is heavily censored.
- MacsHeadroom 3y agoThe official chat finetuned version is censored, the base model is not. The base model is what everyone uses to create their own finetunes, like OpenHermes, Wizard, etc.
- rightbyte 3y agoMalware creation? How is malware distinct from software in general from the LLMs perspective? Like, any software that has some integrated update over the internet feature is potential malware.
- not2b 3y agoThe more interesting security issue, to me, is the LLM analog to cross-site scripting attacks that Simon Willison has written so much about. If we have an LLM based tool that can process text that might come from anywhere and email a summary (meaning that the input might be tainted and it can send email), someone can embed something in the text that the LLM will interpret as a command, which might override the user's intent and send someone else confidential information. We have no analog to quotes, there's one token stream.
- dwaltrip 3y agoCouldn’t we architect or train the models to differentiate between streams of input? It’s a current design choice for all tokens to be the same. Think of humans. Any sensory input we receive is continuously and automatically contextualized alongside all other simultaneous sensory inputs. I don’t consider words spoken to me by person A to be the same as those of person B. I believe there’s a little bit of this already with the system prompt in ChatGPT?
- not2b 3y agoPossibly there's a way to do that. Right now, LLMs aren't architected that way. And no, ChatGPT doesn't do that. The system prompt comes first, hidden from the user and preceding the user input but in the same stream, and there's lots of training and feedback, but all they are doing is making it more difficult for later input to override the system prompt, it's still possible, as has been shown repeatedly.
- dragonwriter 3y ago> Couldn’t we architect or train the models to differentiate between streams of input? Could you? absolutely. Would it solve this problem? Maybe. Would it make training LLMs to do useful tasks much harder and vastly increase the volume of training data necessary? For sure. > I believe there’s a little bit of this already with the system prompt in ChatGPT? Probably not. Likely, both the controllable "system prompt" you can change via the API and probably any hidden system prompt is part of the same prompt as the rest of the prompt, though its deliminited by some token sequence when fed to the model (chat-tuned public LLMs also do this, with different delimiting patterns.)