3 ms·
"Reward hacking" has to be a similar problem space as "sycophancy", no?
by someothherguyy 1y ago
"Reward hacking" has to be a similar problem space as "sycophancy", no?
- cubefox 1y agoSycophancy is one form of RLHF induced reward hacking, but reasoning training (RLVR) can also induce other forms of reward hacking. OpenAIs models are particularly affected. See https://www.lesswrong.com/posts/rKC4xJFkxm6cNq4i9/reward-hacking-is-becoming-more-sophisticated-and-deliberate https://www.lesswrong.com/posts/rKC4xJFkxm6cNq4i9/reward-hac...
- cyanydeez 1y agokeep in mind these models are being taught to talk to each other, so, probably a trick theyre using on each other
- klysm 1y agoReward hacking is literally just overfitting with a different name no?
- n2d4 1y agoThey're different concepts with similar symptoms. Overfitting is when a model doesn't generalize well during training. Reward hacking happens after training, and it's when the model does something that's technically correct but probably not what a human would've done or wanted; like hardcoding fixes for test cases.