22 ms·
You could have spelled out the abbreviation: reinforcement learning with human feedback. Yes, it depends on what the human feedback is and how strongly that af
by xapata 3y ago
You could have spelled out the abbreviation: reinforcement learning with human feedback. Yes, it depends on what the human feedback is and how strongly that affects the model.
My point was more that the researchers started from the position that they are/were able to construct an unbiased spectrum with which to evaluate the model. I am skeptical.