4 ms·
Sure, these particular examples are concerned with morality, but the problem is more general and limits the value of language models because they can be hacked
by yeck 3y ago
Sure, these particular examples are concerned with morality, but the problem is more general and limits the value of language models because they can be hacked for other purposes. A good example that's been going around is having an agential model that manages you emails. Someone sends you an email using prompt injection to compel the agent to delete all your emails. Or forward all you emails to another address.
If there isn't a way to secure the behaviour of AI models against reliable exploits then the utility of the models is dramatically limited.
- TheAceOfHearts 3y agoUse multiple models or gap their capabilities.
- bick_nyers 3y agoDon't allow it to delete emails (maybe allow mark for delete in 30 days?) and have a whitelist of acceptable forwarding addresses or push a confirm/deny notification to a manual reviewer. It's like AI diagnosis, we aren't going to run it full stop automated without safeguards on top or manual review for a long time.