4 ms·
> In my specific case it was the bias to provide an answer to a question That seems to be a reasonably expected result of the "instruction post-training" fine
by PeterisP 2y ago
> In my specific case it was the bias to provide an answer to a question
That seems to be a reasonably expected result of the "instruction post-training" finetuning with RLHF or otherwise. If for some reason you don't want this behavior, you can avoid this by using a model version that just has the core language modeling without that finetuning, e.g. the llama models have such a version available.
- XenophileJKO 2y agoWell in this specific case, the logic I was asking the model to do was. (Highly paraphrased..) 1. Inventory the retrieved items. 2. Determine their relevance. 3. Pick the most relevant or if none of the retrieved items is relevant return an alternative message. What the model will do is add new items into (1) if none of the retrieved items are relevant. If you add some steps between 1 and 2.. it stops doing that.