4 ms·
Do you mean that the gpt creators cannot backtrack an answer to understand how the model came up with it? If it’s such a black box how do they evolve it? Trial
by fruit2020 3y ago
Do you mean that the gpt creators cannot backtrack an answer to understand how the model came up with it? If it’s such a black box how do they evolve it? Trial and error?
- p-e-w 3y agoThey can backtrack it of course, but the result is just billions of numbers – not any sort of "insight".
- HarHarVeryFunny 3y agoAt the end of the day perhaps the most insight we'll get into why the model is saying what it does will be to ask it! Far from ideal of course, and no better than asking a person why they said/did something (which is often an after-the-fact guess). However, at least any such explanation may be using the same internal model/reasoning as what generated the speech in the first place, so conversational probing may support some sort of triangulation into what was behind it!
- iamflimflam1 3y agoTrial and error is pretty much what training is. You feed an input in and use the error to update the network. What is surprising with these models is that the simple training leads to emergent behaviour that is much more powerful than what you’d expect from the training data. With RHLF post training you can tweak these emergent behaviours by having a human (or a model trained to act like a human) give feedback on how good the output is. So far I’ve not seen any good explanations for how this emergent behaviour happens or how it can be reverse engineered.
- simonh 3y agoBasically yes, they use a system called RLHF for Reinforcement Learning with Human Feedback. At a super high level you train your model on source texts. Then you have it generate responses from prompts. Humans rate these responses to select the best ones which updates the model, but you also train a new reward model to mimic the human rankings. Then you train the original model by having it generate millions of responses which are ranked by the rearward model. When I explained this to my brother he literally spat out his tea in horror. This allows you to train at huge scale, many orders of magnitude beyond what you could achieve with just human ranking. The problem is this relies on the reward model accurately capturing what makes a response ‘better’. What it’s actually doing is learning what responses get ranked highly by humans, for whatever reason. Hence the risk of LLMs becoming emotionally manipulative sycophants. It turns out alignment is a really hard problem.
- HarHarVeryFunny 3y agoNeural nets are not generally trained through evolution ("trial and error"), but rather via error minimization, and this is how these GPT models are trained. The basic idea is that the neural net is just a mathematical function, with lots of parameters that control how it calculates it's output, that derives an output value (or set of values) for any input. During training, the neural net also calculates an error (aka "loss") value representing the difference between it's current (at this stage of training) output value and what it was told is the preferred output value for the current input. The process of training is done by slowly adjusting the neural net parameters until these calculated output errors are as small as possible for as many of the training examples as possible. The way these errors are reduced/minimized is by using the derivative (slope) of the neural network function - we want to follow the slope of the error function downhill to a place where the error value is lower, and this is done by adjusting the parameter values using partial derivatives. The details of this downhill slope following (the "backprop" algorithm) are a bit complex, but you can visualize it as a 3-D hilly landscape where the height of the hills represents the size of the error, and the goal it to get into the lowest valley of the landscape (corresponding to the lowest error). If your current lat/long position in the landscape is (x, y) and you know the slope of the hill you are on, then you can move downhill towards the valley by moving a bit in the appropriate direction from (x,y) to (x+dx, y+dy). These x, y values represent the parameters of the network, so by continually tweaking them from (x,y) to (x+dx,y+dy) for each training sample, you are slowly moving down the error hill in the right direction towards the valley of lowest error.
- pmoriarty 3y agoIn other words, we can tell the neutral nets when they're getting "warmer" or "colder" to desirable speech, but we don't know how they do it.
- HarHarVeryFunny 3y agoWell sort of... The odd thing about large transformers is that there is such a huge qualitative difference between what they learn (hence how they behave) and how they are trained, so it's hard to say that this predict-next-word error feedback is directly controlling their inference behavior. Given what the model is learning, it's perhaps best to regard this predict-next-word feedback not as "this is what I'd like you to generate", but rather something more indirect like "learn to generate something like this, and you'll have learnt what I want you to learn". A bit like Karate Kid and "wax on, wax off", perhaps! The actual desirability of what the model is generating, which depends on what you want to use it for, is really controlled by subsequent training steps, such as: 1) Fine tuning for instruction (prompt) following and conversational ability (this is the difference between ChatGPT and the underlying raw GPT-3 model) 2) Goal-based reinforcement learning to stop the model from generating undesirable content such as telling suicidal people to kill themselves, etc, etc.