4 ms·Direct Preference Optimization: Your Language Model Is a Reward Model3 points by ntonozzi 3y ago