4 ms·
Are you trying to argue that RLHF is the same as the algorithm used to pretrain and finetune models, because they are both searches in model space that reward g
by hexaga 3y ago
Are you trying to argue that RLHF is the same as the algorithm used to pretrain and finetune models, because they are both searches in model space that reward good answers and penalize bad answers?
If you only look at the very high level superficial details, they could be said to be similar. That is all. The details differ in nontrivial ways.
> They did not alter the algorithm that generated the model and they also did not directly alter the model itself. The neural net was modified indirectly through training.
It is beyond contesting that multiple, different, algorithms are used to train SOTA language models. RLHF exists because pretraining on a broad corpus for predictive loss minimization, then instruction tuning on specialized datasets, is insufficient in various ways (alignment, "racism", etc).
RLHF and pretraining/finetuning use different algorithms. They are similar in some ways - they can both incorporate backprop with a stateful optimizer like AdamW. This does not make them the same.
RLHF computes loss differently from how it is done in pretraining and finetuning. That, alone, constitutes a modification of the logic behind the weighting process.
Once again, the logic behind the weighting process refers to the training algorithm. Not the transformer architecture. Not the weights themselves. The process that decides what the weights are.
> The training process is a search in model space where bad answers are penalized and good answers are rewarded. The algorithm itself doesn't need altering to find a "less racist" model. There's nothing special about racism that requires a fundamental change in the transformer architecture.
Neither I nor the op claimed that anyone changed the transformer architecture to make it less racist. See my last reply w.r.t. reinterpretation of claims in service of strawman construction.
> They trained the shit out of it by RLHF (Reinforcement Learning from Human Feedback) until it was less racist.
As per above, RLHF is a nontrivially different process to what was used to train the too racist model in the first place.
- xcv123 3y ago> RLHF computes loss differently from how it is done in pretraining and finetuning. That, alone, constitutes a modification of the logic behind the weighting process. You are confusing yourself and still not quite understanding the basic concept. Machine learning has three parts: 1. Training dataset 2. Learning algorithm. 3. The model. The data is input to the learning algorithm which generates the model. To make ChatGPT less "racist" they provided new training data and rewarded it for the correct answers. The process or algorithm of RHLF was not modified specially to handle "racism". These are general learning algorithms. They simply added new training data. This is basic Machine Learning 101.
- hexaga 3y agoAh, I see. You are unfamiliar with how SOTA language models are actually trained in current day, what RLHF is in detail, and how it differs from other training phases. I'd encourage you to read through the GPT3 [0] and llama2 [1] papers for more salient descriptions of the specific mechanics involved and why RLHF is used at all, before trying to argue from some kind of introduction to ml 101 level simplification of how SOTA models are produced. - [0]: https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165 - [1]: https://arxiv.org/abs/2307.09288 https://arxiv.org/abs/2307.09288 For example, from the llama2 paper, section "3.2 Reinforcement Learning with Human Feedback (RLHF)": > RLHF is a model training procedure that is applied to a fine-tuned language model to further align model behavior with human preferences and instruction following. We collect data that represents empirically sampled human preferences, whereby human annotators select which of two model outputs they prefer. This human feedback is subsequently used to train a reward model, which learns patterns in the preferences of the human annotators and can then automate preference decisions. The rest of 3.2 should be enlightening as well (in service to understanding how RLHF is actually quite different from the training algorithms used in pretraining or finetuning). > The data is input to the learning algorithm which generates the model. > To make ChatGPT less "racist" they provided new training data and rewarded it for the correct answers. For frontier language models, this is strictly incorrect. Multiple rounds of different training algorithms are used, each with its own objectives. RLHF is more complex than just 'put better data into same old pretraining algorithm.' > The process or algorithm of RHLF was not modified specially to handle "racism". These are general learning algorithms. They simply added new training data. Nobody is claiming RLHF was modified, but that the 'logic behind the weighting methods' are modified. Pretraining, then finetuning on specialized dataset does not produce a sufficiently aligned (non-racist, gender-neutral, etc) model, no matter how much extra data you throw at it: the answer to this is the introduction of a different algorithm, RLHF, which more accurately fits the model to desired alignment than unsupervised autoregressive loss minimization does. The sequence of events here is: 1. unsupervised pretraining on mixed dataset eval result: "it's just continuing the prompt instead of answering questions or being useful" researcher response: "let's put more focused data in to make it respond better" 2. instruction tuning on task-specific dataset eval result: "it's answering questions now but it has inherited biases from the pretraining corpus (it's too racist, sexist, or otherwise not aligned with desired values)" researcher response: "we've hit the limit on what you can reasonably accomplish with this training algorithm, let's try something else" (see GPT3, section "5. Limitations") 3. rlhf on human preference dataset eval result: "well, it's less racist now" > You are confusing yourself and still not quite understanding the basic concept. Once again, as per my last reply: "If you only look at the very high level superficial details, they could be said to be similar. That is all. The details differ in nontrivial ways." The basic ontology that you're trying to argue exists, is counterfactual. It boils down to: "put data into training algorithm, get model out. if it's too racist, put aligned data in and get non-racist model out." It's not hard to understand, you're just wrong. SOTA models are not produced this way. Full stop. The details of training phases are nontrivially different both algorithmically, and in the datasets used. At this point we've exhausted all reasonable discourse; it's clear that you're not engaging in good faith and I remain convinced that your position falls apart under the most trivial application of rigorous analysis. If you have any actual source to back up your claims, I'd love to see it. Beyond that, good day.