3 ms·
Reinforcement Learning from Human Feedback
https://arxiv.org/abs/2504.12501 https://arxiv.org/abs/2504.12501
- klelatti 8mo agoWeb version with links, etc: https://rlhfbook.com/ https://rlhfbook.com/
- dang 8mo agoThanks! We've switched to that above from https://arxiv.org/abs/2504.12501 https://arxiv.org/abs/2504.12501, and put the latter in the toptext.
- iisweetheartii 8mo ago[dead]
- verdverm 8mo agoLast time I saw Nathan say something about the book, he's actively working on the next version and looking for feedback, check his socials
- leggerss 8mo agoYou could say he's also learning from human feedback
- dang 8mo agoRelated. Others? RLHF Book - https://news.ycombinator.com/item?id=42902936 https://news.ycombinator.com/item?id=42902936 - Feb 2025 (37 comments)