3 ms·Direct Preference Optimization: Your Language Model Is Secretly A Reward Model1 points by optimalsolver 3y ago