4 ms·
Speaking frankly, this seems like a misread to me. Ascribing an author's intent to the output of a program they wrote is not the same as believing they manually
by hexaga 3y ago
Speaking frankly, this seems like a misread to me. Ascribing an author's intent to the output of a program they wrote is not the same as believing they manually performed the actions of the program.
- xcv123 3y agohttps://news.ycombinator.com/item?id=37776217 https://news.ycombinator.com/item?id=37776217
- hexaga 3y ago> [...] modulated the logic behind the Weighting methods. I fail to see how this is not literally the case. The 'logic behind the weighting methods' refers to something like cross-entropy loss minimization via AdamW (or insert optimizer / loss function of choice) over some dataset. I read that as 'the process that decides the weights'. If the model is unaligned ("too racist"), fine tuning with rlhf or alternative of choice is a modulation of the aforementioned 'logic behind the weighting methods'. It is a modification to the process that is deciding weights. Judging by your response, you read 'logic behind the weighting methods' as 'the weights'. I don't think this is a reasonable interpretation unless you're trying to construct a strawman to argue against. Sanity check via chatgpt: > In the context of generative pretrained transformers (GPT), "logic behind the weighting methods" refers to the rationale or reasoning behind the techniques used to assign different weights to words or tokens during the model's training process. - https://chat.openai.com/share/8af4ce45-2c4a-485a-ae8a-e1e016220f90 https://chat.openai.com/share/8af4ce45-2c4a-485a-ae8a-e1e016...
- xcv123 3y ago> Judging by your response, you read 'logic behind the weighting methods' as 'the weights'. I don't think this is a reasonable interpretation unless you're trying to construct a strawman to argue against. So now we are talking about two separate things. The algorithm used to train the model, and the model itself. They did not alter the algorithm that generated the model and they also did not directly alter the model itself. The neural net was modified indirectly through training. That's machine learning 101. When fine tuning a model we do not alter the training algorithm that generates the model. As far as we know, no one is manually tweaking weights or manually editing nodes in its neural network (yet). That is far beyond current knowledge and abilities but is an active field of research in the very early stages. The training process is a search in model space where bad answers are penalized and good answers are rewarded. The algorithm itself doesn't need altering to find a "less racist" model. There's nothing special about racism that requires a fundamental change in the transformer architecture. It's just language and semantics. They trained the shit out of it by RLHF (Reinforcement Learning from Human Feedback) until it was less racist.
- hexaga 3y agoAre you trying to argue that RLHF is the same as the algorithm used to pretrain and finetune models, because they are both searches in model space that reward good answers and penalize bad answers? If you only look at the very high level superficial details, they could be said to be similar. That is all. The details differ in nontrivial ways. > They did not alter the algorithm that generated the model and they also did not directly alter the model itself. The neural net was modified indirectly through training. It is beyond contesting that multiple, different, algorithms are used to train SOTA language models. RLHF exists because pretraining on a broad corpus for predictive loss minimization, then instruction tuning on specialized datasets, is insufficient in various ways (alignment, "racism", etc). RLHF and pretraining/finetuning use different algorithms. They are similar in some ways - they can both incorporate backprop with a stateful optimizer like AdamW. This does not make them the same. RLHF computes loss differently from how it is done in pretraining and finetuning. That, alone, constitutes a modification of the logic behind the weighting process. Once again, the logic behind the weighting process refers to the training algorithm. Not the transformer architecture. Not the weights themselves. The process that decides what the weights are. > The training process is a search in model space where bad answers are penalized and good answers are rewarded. The algorithm itself doesn't need altering to find a "less racist" model. There's nothing special about racism that requires a fundamental change in the transformer architecture. Neither I nor the op claimed that anyone changed the transformer architecture to make it less racist. See my last reply w.r.t. reinterpretation of claims in service of strawman construction. > They trained the shit out of it by RLHF (Reinforcement Learning from Human Feedback) until it was less racist. As per above, RLHF is a nontrivially different process to what was used to train the too racist model in the first place.
- xcv123 3y ago> RLHF computes loss differently from how it is done in pretraining and finetuning. That, alone, constitutes a modification of the logic behind the weighting process. You are confusing yourself and still not quite understanding the basic concept. Machine learning has three parts: 1. Training dataset 2. Learning algorithm. 3. The model. The data is input to the learning algorithm which generates the model. To make ChatGPT less "racist" they provided new training data and rewarded it for the correct answers. The process or algorithm of RHLF was not modified specially to handle "racism". These are general learning algorithms. They simply added new training data. This is basic Machine Learning 101.