3 ms·
It is fine tuned to maximize reward though, not likelihood. And it provides an answer in both cases, just not as well.
by jumpCastle 3y ago
It is fine tuned to maximize reward though, not likelihood. And it provides an answer in both cases, just not as well.
- stoniejohnson 3y agoSo since a model is fine tuned via RLHF my point doesn't stand? Genuine question; it would be interesting if some other mechanism was at play here.
- jumpCastle 3y agoFor an answer I would expect it to get the same reward for both question orderings. So naively I would expect it to not be affected by the ordering.