3 ms·
>These include reward hacking which is a huge problem with LLMs and why they lie and hallucinate so persistently, power seeking behaviour, and developing unanti
by yanderekko 3y ago
>These include reward hacking which is a huge problem with LLMs and why they lie and hallucinate so persistently, power seeking behaviour, and developing unanticipated harmful instrumental goals.
I'm not sure I'd call it reward hacking when everyone understands that LLMs are just trying to engage in text prediction and that this can often be a poor proxy for trying to generate "useful" text. A paperclip maximizer that ends up turning the world into paperclips rather than creating a new cool paperclip factory has an alignment problem, but it did not reward hack.
- simonh 3y agoWhat the paper clip maximiser is doing is generating harmful unintended instrumental goals, such as killing the humans so that they can’t stop it making more Paperclips.
- yanderekko 3y agoSure, but I wouldn't consider it reward hacking if we knowingly gave it the goal of maximizing papercips and it did this. Maybe I'm being pedantic, but that would be a pretty clear due diligence failure on the prompter's fault.