4 ms·
Also it is important to note that GPT shows large miss-alignments. The problem comes from the fact, that it is hard or impossible to give an objective what GPT
by fxj 3y ago
Also it is important to note that GPT shows large miss-alignments. The problem comes from the fact, that it is hard or impossible to give an objective what GPT should be optimized to (nobody knows the truth), so proxys are used. One proxy is that it should make the user happy and the user should give many thumbs up. But this does not mean that it has to give the "correct" answer, which the user himself might not know in the first place. So it invents things because during the reinforcement learning users were happy with these answers. A funny example is the github co-pilot which writes buggy code, because it thinks this is what the user wants. Here is a video about that:
https://youtu.be/viJt_DXTfwA?t=1767 https://youtu.be/viJt_DXTfwA?t=1767