3 ms·
To save others from searching, RLHF [1] is Reinforcement Learning from Human Feedback [1] https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedb
by iroddis 3y ago
To save others from searching, RLHF [1] is Reinforcement Learning from Human Feedback
[1] https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback?wprov=sfti1 https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu...
- belter 3y agoAnd to put it in context...RLHF is the special sauce behind ChatGPT