4 ms·
Ah, I see. You are unfamiliar with how SOTA language models are actually trained in current day, what RLHF is in detail, and how it differs from other training
by hexaga 3y ago
Ah, I see. You are unfamiliar with how SOTA language models are actually trained in current day, what RLHF is in detail, and how it differs from other training phases. I'd encourage you to read through the GPT3 [0] and llama2 [1] papers for more salient descriptions of the specific mechanics involved and why RLHF is used at all, before trying to argue from some kind of introduction to ml 101 level simplification of how SOTA models are produced.
- [0]: https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165
- [1]: https://arxiv.org/abs/2307.09288 https://arxiv.org/abs/2307.09288
For example, from the llama2 paper, section "3.2 Reinforcement Learning with Human Feedback (RLHF)":
> RLHF is a model training procedure that is applied to a fine-tuned language model to further align model behavior with human preferences and instruction following. We collect data that represents empirically sampled human preferences, whereby human annotators select which of two model outputs they prefer. This human feedback is subsequently used to train a reward model, which learns patterns in the preferences of the human annotators and can then automate preference decisions.
The rest of 3.2 should be enlightening as well (in service to understanding how RLHF is actually quite different from the training algorithms used in pretraining or finetuning).
> The data is input to the learning algorithm which generates the model.
> To make ChatGPT less "racist" they provided new training data and rewarded it for the correct answers.
For frontier language models, this is strictly incorrect. Multiple rounds of different training algorithms are used, each with its own objectives. RLHF is more complex than just 'put better data into same old pretraining algorithm.'
> The process or algorithm of RHLF was not modified specially to handle "racism". These are general learning algorithms. They simply added new training data.
Nobody is claiming RLHF was modified, but that the 'logic behind the weighting methods' are modified. Pretraining, then finetuning on specialized dataset does not produce a sufficiently aligned (non-racist, gender-neutral, etc) model, no matter how much extra data you throw at it: the answer to this is the introduction of a different algorithm, RLHF, which more accurately fits the model to desired alignment than unsupervised autoregressive loss minimization does.
The sequence of events here is:
1. unsupervised pretraining on mixed dataset
eval result: "it's just continuing the prompt instead of answering questions or being useful"
researcher response: "let's put more focused data in to make it respond better"
2. instruction tuning on task-specific dataset
eval result: "it's answering questions now but it has inherited biases from the pretraining corpus (it's too racist, sexist, or otherwise not aligned with desired values)"
researcher response: "we've hit the limit on what you can reasonably accomplish with this training algorithm, let's try something else" (see GPT3, section "5. Limitations")
3. rlhf on human preference dataset
eval result: "well, it's less racist now"
> You are confusing yourself and still not quite understanding the basic concept.
Once again, as per my last reply: "If you only look at the very high level superficial details, they could be said to be similar. That is all. The details differ in nontrivial ways."
The basic ontology that you're trying to argue exists, is counterfactual. It boils down to: "put data into training algorithm, get model out. if it's too racist, put aligned data in and get non-racist model out." It's not hard to understand, you're just wrong. SOTA models are not produced this way. Full stop. The details of training phases are nontrivially different both algorithmically, and in the datasets used.
At this point we've exhausted all reasonable discourse; it's clear that you're not engaging in good faith and I remain convinced that your position falls apart under the most trivial application of rigorous analysis. If you have any actual source to back up your claims, I'd love to see it. Beyond that, good day.
- xcv123 3y ago> the answer to this is the introduction of a different algorithm, RLHF, which more accurately fits the model to desired alignment than unsupervised autoregressive loss minimization does. FFS. That's exactly what I have been trying to explain to you. RHLF is the general method. So you agree with me. I am fully aware that there are different algorithms used at different stages of the training process. No one is hardcoding racism out of the model or hardcoding antiracism in the training algorithm. That is learned automatically. It's still the same basic machine learning 101 concept. > At this point we've exhausted all reasonable discourse; it's clear that you're not engaging in good faith and I remain convinced that your position falls apart under the most trivial application of rigorous analysis. If you have any actual source to back up your claims, I'd love to see it. Beyond that, good day. Totally unnecessary and uncalled for. You have repeatedly accused me of bad faith which lowers the discussion to a childish level of vindictive garbage. You were not making your point clear before and I was making the effort to respond sincerely to that, trying explain what I thought you misunderstood. It seems that you are still confused and are arguing my point for me without realizing it. Thanks for making my point clearer and adding interesting technical references to back it up. You have provided the sources to back up my claims. Cheers. Have a good one.
- hexaga 3y ago> [...] As per my last reply, I don't believe further debate of this point would be productive. We've both made our arguments. I'm comfortable letting mine stand on their own merits. If you'd like to start a meta-discussion about this discussion, and any grievances therein, I'll happily participate if we can keep it civil. I think understanding why and how debates can breakdown at the limits (and what those limits are) provides real value in avoiding such breakdowns in the future. > Thanks for making my point clearer and adding interesting technical references to back it up. You have provided the sources to back up my claims. I am glad to have made a positive impact on you in some small way, at least. > Totally unnecessary and uncalled for. You have repeatedly accused me of bad faith which lowers the discussion to a childish level of vindictive garbage. You were not making your point clear before and I was making the effort to respond sincerely to that, trying explain what I thought you misunderstood. It seems that you are still confused and are arguing my point for me without realizing it. In the spirit of the above, if you are being sincere: I must genuinely urge you to consider why someone might believe your arguments to be unsatisfactory, and how charitable it is not to only consider incoherent formulations of the arguments of others. Put simply: if you see someone's argument as trivially unsound, is it possible you have misunderstood their position? If you're admitting that you did not understand my argument with sufficient clarity to contest it, on what grounds can you claim to know what I've misunderstood in the making of it? My claim follows from this line of reasoning. You are, necessarily, by your own admission, arguing against a strawman.