4 ms·
I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.
by visarga 3mo ago
I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.
- redanddead 3mo agoWhat makes you say that
- versteegen 3mo agoMore accurate to say RLHF aligns models to human preferences, most significantly to be helpful.