3 ms·
I don't think it does. And there is a pretty big risk that you end up picking up on some quirk ("bias") of your reward model that doesn't reflect reality -- GPT
by huac 3y ago
I don't think it does. And there is a pretty big risk that you end up picking up on some quirk ("bias") of your reward model that doesn't reflect reality -- GPT4 preferring longer answers is one such commonly observed bias. AFAIK there is not a great theoretical basis for why we can avoid mode collapse, except empirically the models are good enough to survive some bootstrapping.