5 ms·
I don't follow. Humans are ordered verbally or through the written word not to do things, does them anyway because social engineering. LLMs are have guardrail
by TerrifiedMouse 3y ago
I don't follow.
Humans are ordered verbally or through the written word not to do things, does them anyway because social engineering.
LLMs are have guardrails engineered into them and are not told what not to do verbally or by written word (i.e. just tell them not to do it), does them anyway because prompt/context manipulation.
I'm not criticizing the failure of the LLM to follow orders. I'm criticizing the way orders have to be given.
- bastawhiz 3y agoFine tuning doesn't help avoid jailbreaking, it just makes it harder. So no, you're not always mucking with prompts and contexts. LLMs fail at following orders in almost exactly the same ways that humans do, much to everyone's chagrin.
- famouswaffles 3y ago>LLMs are have guardrails engineered into them They don't. What do people think LLMs are lol ? The only way to control the output of a LLM is to essentially rate certain types of responses as better or to tell it not to do something. any other "guardrails" are outside direct influence of the LLM (i.e a separate classifier that blocks certain words). Nobody is "engineering" anything into LLMs.
- TerrifiedMouse 3y ago> The only way to control the output of a LLM is to essentially rate certain types of responses as better Which is my point. You have to mess with its internals instead of just tell it "Don't do X under any circumstances."
- famouswaffles 3y agoFirst of all, no you don't have to. Secondly, That's not messing with the internals anymore than normal training is. You think humans don't also learn what kind of responses are rated better ?
- deleted 3y ago[deleted]
- squeaky-clean 3y agoColloquially, LLM tends to refer to the entire product. Not just the model weights. For example technically GPT-4 isn't an LLM, it's 16 LLMs in a trench coat.