3 ms·
apparently (according to the blog post) that's a result of the RL human preference fine-tuning - the human rankers preferred longer more in-depth answers
by make3 4y ago
apparently (according to the blog post) that's a result of the RL human preference fine-tuning - the human rankers preferred longer more in-depth answers
- TheCaptain4815 4y agoWasn't the instruct model created using that same strategy?